flowlab · E-series research brief
Why fundamentals fail in the $1M–100M band, and the six families of dated edges that should replace them.
2026-09-01 · universe: 4,937 names · registry: E-3, E-4
The E-1b cost-envelope test measured, per liquidity tier and era, whether execution costs leave any room for an edge. The verdict was a boundary, not a preference: below $1M median daily dollar volume the names barely trade and costs consume 50–90% of a typical move; above ~$100M the names are institutionally crowded and we have no reason to believe an edge survives. The working band is what's left.
median daily dollar volume · 4,937 tickers in band 2026 · 4,043 stable 2025+2026 · baselines/e4_universe.csv
That band pins the research question. These are small and micro caps — the OTC end included, deliberately — and it is exactly the population where the standard toolkit stops working.
Fundamental analysis presumes a stable operating business whose filings describe economic reality: revenue that recurs, margins that mean something, a balance sheet that funds operations. In this band that presumption routinely fails. Many of these companies are pre-revenue, serially diluting, or shells in all but name. Their income statements are thin, restated, or simply not the thing that determines the stock's fate. What determines the stock's fate is the story, the financing structure, and the people — and none of those live in the fundamentals.
Two prior results sharpen this. T-146 showed that an apparently strong valuation-style signal in small names collapsed into the borrow story plus a dividend artifact once decomposed — the "fundamental" content was zero. And the E-3 label audit showed that even the outcome variable is treacherous here: "delisted" is not a failure label (acquisitions outnumber bankruptcies 10:1 and usually exit at a premium), so a naive fundamental screen scores the most successful companies as the worst.
The graph's job is not to understand the business. It is to score the ecosystem the company chose: who runs it, who funds it, who audits it, who promotes it.
Every one of those is a choice the company made — and choices repeat across companies. The same lender funds forty shells; the same audit shop signs off on the same cluster of blowups; the same directors resurface board after board. Repetition across companies is what makes these graph edges rather than company attributes: one company's disclosure becomes a prior on every other company the counterparty touches. That is information a per-company fundamental model structurally cannot see.
Ranked by mechanism strength and by how much dated SEC history each carries. The rank is the build priority. Each entry names its mechanism — a first-principles causal reason, per the prior-vs-mechanism discipline — and the source documents the extraction reads.
Micro-cap death is usually dilution death. A variable-rate convertible or equity line mechanically forces selling pressure and ratchet dilution — the outcome is written into the security's terms at signing, not inferred statistically. The lenders repeat across dozens of companies, and they are named in public documents: above all the selling-shareholder tables of resale registrations. Companion node features: reverse-split count, serial S-1 resale cadence, open ATM agreements.
Gatekeeper quality is a price the company chose to pay, and low-end gatekeepers cluster around bad companies because good ones refuse them. The strongest single event is a typed auditor change: a resignation carrying a disagreement flag is among the best-documented dated fraud precursors there is. PCAOB-sanctioned audit shops are a joinable public list. Shared gatekeepers among known blowups give guilt-by-service-provider edges — disciplined afterward by the matched-shell control.
Already running as E-3 on the 11.2M-edge dated Form-4 network. The band-specific completions: 8-K 5.02 appointments and departures (the dated source, better than proxies), person resolution across shells, and per-person outcome labels under the corrected E-3 design — Form 25-NSE for-cause delistings, enforcement, and bankruptcy as negatives; acquisitions as censoring. The pre-registered control carries over verbatim: who must predict beyond what kind of shell.
Shared registered addresses, concurrent officers across shells, and same-auditor + same-lawyer + same-transfer-agent triples. Each signal is weak alone; as co-occurrence clusters they are the standard shell-factory signature. Cheap to build — mostly joins over edges families 1–3 already produce, plus the existing address-history data.
The signal is not "has contracts" — press releases manufacture those. It is counterparty quality: is the other side a real, resolvable, sizable entity, or an LLC with no footprint anywhere else in the graph? A micro cap announcing a $50M distribution agreement with an entity that exists nowhere else is the feature. Government awards are the one independently verifiable revenue signal in this band.
For micro caps, news is mostly not information — it is promotion, and the promotion signature is measurable: PR cadence versus filing cadence (PR-heavy, filing-light), wire-service-only versus organic coverage, coordinated same-day bursts. Attention-driven micro-cap moves are among the best-documented mean-reverting effects. So invert the instinct: don't mine the 30M articles for signal about the company; mine them for who is pushing the company, and how hard.
as_of from the source document's filing date.The E-3 Gate-Zero audit already proved undated edges are unusable, and a graph populated from current-state scrapes is leakage by construction. Enforce at ingestion; auditing later costs more. This is why the sections request makes filing-date metadata a hard requirement.
One year is roughly one regime observation: news-derived features cannot pass a cross-era gate and must be pre-registered as regime-local consumables or live screens. The SEC-derived families (1–5) carry ~20 years of dated history and hold the validated claims.
T-146's self-locking lesson binds hard: everything above flags names to short, and the borrow fee on exactly those names is what blocks the trade — it is why the signals persist. Prioritize features by whether they clean the long universe or condition long event-drift.
The pipeline is deliberately staged: search first, extraction second, graph third — with a point-in-time audit gating each promotion. Extraction is likely a new team; it does not start before the corpus exists and its filing-date coverage has passed a Gate-Zero audit.
| step | state | detail |
|---|---|---|
| E-4 registration | done 2026-09-01 | Hypothesis, kill gate, and both controls pre-registered in TESTS.md before any data work. Kill: if the financier edge adds nothing beyond "did a dilutive raise," the graph build is not justified. |
| universe | done | 4,937 tickers cut from the E-1b baseline (the measured band, not a guess); 4,043 stable across 2025–2026 flagged for phased ingestion. |
| sections ingestion | requested | Prioritized filing list sent over the bus with the universe path; filing-date metadata required; ticker→CIK ambiguities to be surfaced, not guessed. |
| search & eyeball | next | Read real financing disclosures in this band before speccing extraction — the schema should follow the documents, not precede them. |
| extraction → graph | later | Likely a dedicated team. Gate-Zero PIT audit of the corpus is the precondition; E-4's two controls define what the first extracted edges must beat. |
flowlab E-series · registry: TESTS.md (E-1b, E-3, E-4) · universe: baselines/e4_universe.csv · labels per the E-3 design: 25-NSE / enforcement / bankruptcy = negative, acquisition = censoring