Strategy · Business Information · 2026

Shock
Wave

The incumbents sell public data back to the market at license prices. We take the entire public-derived layer — on free data, with an AI-native entity graph — and leave only the true moat standing.

The thesis

Strip a Bloomberg, FactSet, Capital IQ or Morningstar down to its content and most of it is SEC and regulatory filings, re-keyed and rented back — ownership, insiders, M&A, fundamentals, people, boards, comp, governance, events. That layer isn't a moat. It's extraction. We already run it, and we can take the rest of it for the cost of engineering, not a license.

We mapped 153 companies across the industry and sorted every data domain into four tiers. The point isn't to copy a terminal — it's to see exactly how far free data reaches, and where the real wall is.

TIER 0
HAVE
Already extracted
Insider, XBRL, agreements, comp, 990s, ADV, debt, the graph.
TIER 1
TAKE
Free public data + build
Theirs today, ours with extraction. The attack surface.
TIER 2
BUY
Licensed, not buildable
Pricing, estimates, premium newswire. Buy only if a product needs it.
TIER 3
MOAT
Proprietary
Ratings opinions, expert calls, alt-data panels, EMR. We leave it.

The attack surface

153 mapped. The public layer is wide open.

24
domains at Tier 0–1 — the public-derived layer we take on free data. The roadmap.
38
at Tier 2 — a license wall. Buy only where a product truly needs it.
90
at Tier 3 — genuine proprietary moat. We don't pretend to take it.

Not vaporware — already live

The engine is running.

This isn't a pitch deck for data we wish we had. The Tier-0 layer is in production today, entity-resolved on one graph, with an AI chat that answers grounded and cited.

9.7M
SEC exhibits, full-text searchable
30.8M
clause-level docs, semantic
11.5M
insider transactions, with $ amounts
305K
insider / executive / director profiles
468K
classified agreements, 99% verbatim-quoted
28M
nonprofit board-officer records (990s)
~30M
news articles, entity-linked
1
graph — deals, debt, people, comp, all on CIK/LEI

The next wave

What we take next — on free data.

Each: a dataset a competitor sells whose raw material is free or public. The work is extraction + linkage + history, not a license.
01
International filings — SEDAR+, Companies House, EDINET
in-build · vs 2iQ, Quartr, pfactorial, OpenCorporates
high
02
OCR of scanned EDGAR exhibits
image-only filings → the clause index · the pfactorial hole
high
03
US court dockets & opinions
CourtListener / RECAP · vs Bloomberg Law, Westlaw, Lexis
high
04
Federal contracts & grants
USAspending / SAM.gov · cheap early win
low eff
05
Lobbying & political money
Senate/House LDA, FEC, PTRs · donor affinity
low eff
06
Fund holdings (N-PORT) + proxy votes (N-PX)
EDGAR · the missing half of the US ownership book
med-high
07
Non-GAAP & segment KPIs past XBRL
8-Ks, IR decks · vs Daloopa, Fiscal.ai
high
08
County property records for donor wealth
assessor/recorder sites · the last public asset in the capacity number
high
09
Credit-rating actions from filings & press
8-K items · the action, not the opinion · cheap
low eff
10
Grant graph on 990s · Form D tape · supplier mentions
IRS 990 / EDGAR · coverage + entity resolution
med

Full ranked backlog: 19 Tier-1 datasets, each with its free source and the exact point where the license wall begins.

The honesty that makes it credible

We don't take the moat.

A disruptor that claims everything is noise. We name what we won't chase — and that discipline is the strategy. We take the ~80% of the value that's public-derived at near-zero marginal data cost, and we don't burn capital fighting for the genuinely proprietary last mile.

The wall we leave standing

real-time pricingconsensus estimatesrating opinions ESG scoresexpert-call librariescard / location panels app telemetrysatellite / AIS cargoEMR & gratitude (GOBEL) consumer net-worth modelscontributory people graphspremium newswire

The disruption

Their model is a data license — recurring cost, per seat, per feed. Ours is extraction + an entity graph + AI agents: build once, near-zero marginal data cost, and every new free source compounds onto the same graph.

Match the public-derived layer that is most of what they sell, at a fraction of the price, and the pricing power on that layer collapses. That's the shock wave.

The filings were always public.
We just made them answerable.

Grounded in live corpora (sec_exhibits_* / sec_clauses_* on OpenSearch; insider.transactions; mars agreements, 990s, news) and a 153-company industry tiering, 2026-10-07. Tier counts from the landscape analysis; "do not queue" Tier-2/3 list is deliberate scope, not omission.