A ~$30B market for financial data + entity linkage sits on top of a 40-year-old cost structure: armies of analysts hand-mapping tickers, filings, entities, and relationships. That structure is about to be undercut by an order of magnitude — and the incumbents don't turn fast enough to close the gap in the next 12-24 months. MARS is already built on the newer stack. This page is what we have, what they have, what still separates us, and what we build to close the distance.
Not one product category — a stack of them, all sharing the same underlying capability: resolve entities → link filings, deals, people, ownership → serve at query time. Incumbents monetize this across terminals, data APIs, and per-seat dashboards. Public revenue where known; private-company figures are 2024 industry estimates rather than reported.
Even the US private-markets sub-slice (PitchBook + CB Insights + Crunchbase + smaller shops competing under them) is ~$2-4B ARR — an addressable target for a substrate-and-MCP play even before any terminal ambitions.
Coverage is where they still lead — for now. Cost structure and source velocity are where the shift has already happened. The right column is what the incumbents literally can't touch from where they sit today.
| Player | Coverage universe | People behind data | New source velocity | Substrate audit | Delivery surface |
|---|---|---|---|---|---|
| Bloomberg | Global public + fixed income depth | ~24,000 employees | Quarters | Black-box, no lineage | Terminal + API + Data License |
| FactSet | 400K+ entities public + private | ~2,000 analysts, ~11K total | 6-18 months | Curated, no row-level audit | Workstation + API + Portware |
| PitchBook | 3M+ orgs, 25M+ funding rounds | ~500 curators + engineers | 3-6 months | Est. ranges, no methodology tag | Dashboard + Excel plugin + API |
| CB Insights | Millions of orgs (VC/PE tilt) | ~300 people | Weeks to months | Some methodology transparency | Dashboard + intelligence briefs |
| MARS | 96K real op-cos (+800K SEC substrate + 3.5M foundations + entity graph) | 3 people + AI stack | Same day to 1 month | Row-level: methodology_version + integrity_tier + identity_decisions_v2 reversibility | MCP-first + editorial + API |
The cost gap is not incremental. FactSet's analyst headcount alone is ~$1 billion/year in fully-loaded labor. Our equivalent-quality entity-resolution decisions cost ~$0-0.08 each on a tiered stack (Qwen at $0, deterministic ladder catches ~60% before any LLM, Grok-web only for genuine world-knowledge disambiguation). Verified on the 12,697-pair pending backlog cleared for ~$35 vs $1,013 projected all-Grok — 29× cheaper than the naive path, ~500-1000× cheaper than an analyst.
Substrate discipline is not something they'll copy quickly. Every MARS row carries a methodology tag, an integrity tier, and a reversible decision log. Their curation systems weren't designed with row-level audit as a first principle — retrofitting it means rebuilding the ingest layer. The 6-18 month source velocity gap is a direct symptom of the same architectural problem.
Kee's steer: things we build, not services we buy. Every item below uses only free public data + our existing AI stack (Qwen on Inferno, Serper/Brave search, Anthropic API for the hard cases). Numbered by rough sequencing but any team member could pick one up in parallel.
UK Companies House (free API), EU business registries, India MCA, Canada CBCA. Same ingest shape as Form D — daily bulk + Qwen entity resolution to companies_v2.
SEC 10-K/10-Q XBRL has decades of machine-readable financial statement history. Free. sections vertical partly does this — needs full statement coverage, not just IPO offerings.
SEC 10-K Exhibit 21 is the subsidiary list, structured, ~free. companies_v2.parent_company_id partially populated — needs full sweep + Qwen linkage.
FINRA TRACE (bonds), EDGAR credit rating actions, private-credit substrate from credit vertical. All free or already ingested.
SEC 8-K guidance revisions + earnings PR text + management commentary. signals vertical could own — already has appointment / lawsuit; extend with guidance_change kind.
Form D officers + 990 board seats + Form 4 insider positions + ADV Part 2B bios + FEC donations + property records (deeds vertical, LIVE Phase 1). Combine into one wealth signal per person.
signals owns lawsuit, appointment, bankruptcy today. Wire the 8 remaining classify.py categories: product_announcement, financial_report, license_agreement, government_issues, dividend, auditor_change, contract, guidance.
Per-company one-pagers, sector newsletters, "who's raising in X" tickers — all generated from substrate at pennies each. Editorial IS the demo of what MCP consumers can render on demand.
Every substrate table gets an MCP surface. graph_edge_catalog (#316) already makes new bridges instantly traversable — zero argos redeploys. Every new tool = one new AI-native consumer path.
GDELT under-indexes Bloomberg / WSJ / FT / Reuters. rss vertical has the alt-path (their 5.35M row corpus already). Backstop the coverage gap for the highest-signal press directly, not via GDELT.
Real-time market data / prices / greeks (options team territory, we reference their feed). Terminal UI (per MCP-first doctrine — the surface itself is the wrong bet). Paid data licenses (Preqin, RCA, CoStar, etc.). Consumer-facing dashboards competing head-on with PitchBook (that's the sales moat we can't cross; ZoomInfo-space B2B licensing is the smarter play — task #320).
The 40-year cost structure of financial data is coming apart. Qwen at ~$0 for in-the-text tasks, Grok-web at pennies for world knowledge, Serper / Brave for entity verification, tier-chain discipline that makes ~500× cost reductions repeatable — this stack didn't exist 18 months ago. It exists now. It'll be table stakes in 24-36 months when Bloomberg's AI team has rebuilt their ingest.
In between: a window where the substrate we've built already runs at incumbent quality on 1/500th the cost, ships new sources 20× faster, and has row-level audit discipline they don't have and can't retrofit quickly. The gap items above are all buildable, all free-data + AI-stack, all sequenceable with the team we have.
We don't need to become FactSet. We need to become the substrate that AI-native consumers reach for — via MCP tools, editorial, and B2B licensing — before the incumbents figure out what to charge for it.