Sections — what you can ask
Ten years of filings, indexed by meaning.
The queries the corpus was built for — from tracking one company's disclosures through time, to finding cross-market patterns, to composing the answer with an LLM sitting on top.
Opening beat
"What actually changed in NVIDIA's most recent 10-K compared to the year before?"
retrievesections__filing_change(cik=1045810, form="10-K") returns a per-sub-section diff between NVIDIA's 2026 Item 1A and its 2025 baseline — 3 sub-sections added, 3 modified, 0 removed. Total-change magnitude 0.024 (light rewrite; NVIDIA's Risk Factors don't churn).
notable adda new sub-section on counterparty concentration risk, magnitude 987 on the raw scale — a first-time disclosure that this specific risk is now big enough to warrant its own bolded heading. Body text follows.
NVIDIA quietly added a counterparty-concentration risk sub-section this year that wasn't in the 2025 10-K. It reads as a warning about revenue exposure to a small number of large customers — a new admission for a company that used to disclose customer concentration only obliquely inside "Business."
Two of the three modified sub-sections tightened language around export-control risk. The third softened the cybersecurity-incident language slightly (removed "material breach" as a hypothetical, kept the framework).
01 — one filer, over time
Longitudinal: watch a company's disclosures evolve.
The clean version of what analysts already do by hand — only faster, and on any filer, across the whole 10-year corpus.
How has JPMorgan's Risk Factors section evolved since 2020?
Retrieve
Loop sections__filing_change across 6 years of 10-Ks for cik 19617. Each year returns added / removed / modified sub-sections + magnitude score.
Compose
A six-year narrative of JPMorgan's evolving concerns — rate risk framing shifted in 2022, crypto-exposure disclosure appeared in 2023, an AI-and-model-risk sub-section appeared in 2024. Magnitudes let the agent flag which years mattered.
real datumJPM 2026 vs 2025: mag 0.16 (moderate rewrite — between "cosmetic" and "substantial")
When did AI / ML risk first appear in Apple's 10-K?
Retrieve
Scan each year's Apple Item 1A sub-section labels. First year where topics includes "AI and Machine Learning Risks" is the answer. Then pull that sub-section's body for the actual disclosure text.
Compose
"Apple first added AI/ML risk language in its 2024 10-K, filed October 2024. The initial sub-section framed it primarily as regulatory and reputational; the 2025 revision expanded to include competition from third-party models."
Track supply-chain disclosure in TSMC's 20-F from 2018 through today.
Retrieve
Semantic query "supply chain concentration" on TSMC's 20-F Item 3 sub-sections, one per year. Pull matching sub-sections + surrounding topic labels.
Compose
Not just "did they talk about it" but how the framing evolved — from generic capacity risk in 2018, to explicit geopolitical framing post-2020, to specific customer-tier concentration in the most recent filings.
02 — many filers, one moment
Cross-sectional panels: who else is saying this?
The pre-classified labels turn "does this company disclose X" into a single filter query. The answer is a cohort, not a single filing.
Which 2026 10-Ks disclose remediation of a prior material weakness?
Retrieve
sections__topic_panel(form="10-K", item="9A", label="Remediation of Material Weakness", year=2026) — single filter query.
Compose
Ranked list of filers plus each one's remediation narrative. Immediate short-list for anyone tracking ICFR quality.
real cohortAXON, Xponential Fitness, Ambiq Micro, Seritage Growth Properties, Redwood Mortgage Investors IX, others — 9 filers surfaced in the initial demo
Show me the pharma cohort exposed to Paragraph IV / Hatch-Waxman litigation.
Retrieve
Filter Item 3 (Legal Proceedings) section-grain points where primary_label = "Pharma Paragraph IV / Hatch-Waxman". Sub-cohort by SIC prefix if desired.
Compose
A pharma-litigation prospect list. This label was empirically distinct enough from generic IP-litigation boilerplate to justify its own category during 2026-06 taxonomy refinement.
Every company that flagged a material cybersecurity incident under Item 1C in the last 12 months.
Retrieve
Item 1C section-level points with the "Material Cybersecurity Incident" label + filing_date filter. ~500 filings/year across the corpus (about 8% prevalence).
Compose
Base rate context + the individual disclosure narratives. Analyst gets a "who and what" table without reading every 10-K.
03 — who else looks like this
Peer neighborhoods: who reads like whom?
Because sub-sections are E5-embedded, "similar to" is a real cosine query — not a manually-curated peer set.
Which banks have Risk Factors most similar to JPMorgan's?
Retrieve
sections__peer_similarity(cik=19617, form="10-K", item="1A") — pure semantic overlap across every filer's Item 1A sub-sections.
Compose
A peer set surfaced by content, not by sector code. In demo: Bank of America, Regions, PNC, NBT, American Express, Schwab. No SIC filter used.
How does Uber's MD&A on cost structure compare to Lyft's?
Retrieve
Both filers' Item 7 sub-sections filtered to primary_label ∈ {"Cost Structure", "Operating Expenses"}. Return matched pairs.
Compose
Side-by-side of what each says about the same cost line, in their own words. The LLM can synthesize where they converge, diverge, and which one gives more granularity.
Find clinical-stage biotechs with risk profiles similar to Moderna's late-stage cohort.
Retrieve
Peer-similarity against Moderna's Item 1A, then filter to SIC 2836 (Pharmaceutical Preparations) and topic label "Clinical Development Risks".
Compose
A screen of biotech companies talking about the same trial-stage risks with the same specificity. Useful for peer-set construction that's grounded in actual disclosure, not just market cap or headcount.
04 — change at market scale
Cross-sectional churn: who rewrote what?
This is the Lazy Prices signal at per-filer grain. The academic literature (Cohen-Malloy-Nguyen) argues that YoY 10-K language changes predict returns because the change itself carries information the market hasn't fully priced.
Which S&P 500 companies rewrote their Risk Factors most between 2025 and 2026?
Retrieve
sections__filing_change_screen(form="10-K", item="1A", year=2026, compare_year=2025) filtered by a cik universe.
Compose
Ranked list, magnitude + brief narrative of what changed. An analyst opens the top 10 immediately rather than reading 500 10-Ks.
panel scale10-K Item 1A 2026 vs 2025: 4,233 paired filers, mean magnitude 0.17, top names include Blackstone Mortgage Trust (155 sub-sections added — commercial-real-estate distress narrative), Trump Media post-transformation, and EQT Infrastructure (newly public)
Any 10-K in the last quarter where Legal Proceedings went from boilerplate to substantive?
Retrieve
Item 3 (section-grain) filing_change with a filter for base-rate deviation — specifically transitions from label "No Material Proceedings" to any substantive litigation label.
Compose
Every filer that added their first meaningful Legal Proceedings sub-section this year — likely names for near-term catalyst research.
real panelItem 3 2026 vs 2025: 87% of filers UNCHANGED (boilerplate is stable). The 32-filer / year "substantial-to-major" cohort at the top of the panel is the real signal — Supernus Pharma (Hatch-Waxman), iAnthus Capital, Ready Capital (CRE distress), Angel Studios (IP).
Companies that added going-concern language to their Risk Factors this year.
Retrieve
Semantic search for "going concern substantial doubt" against Item 1A across the year, cross-referenced against absence in the same filer's prior-year Item 1A. Delta = new adoptions.
Compose
Immediate short-list for distressed-credit or short-selling research. Going concern is a strong signal; new-this-year going concern is a stronger one.
06 — new issuers
IPO / prospectus: a whole company from zero.
424B4 and S-1 fill the "no 10-K history yet" gap. Combined with catalyst's lockup calendar, agents can build a complete new-issuer profile at pricing.
For SpaceX's 424B4, give me the offering terms, lockup structure, and top 5 risk factors.
Retrieve
sections__ipo_offering(cik=1181412) for terms + sections__ipo_risk_factors for classified risks. Lockup fields JOIN from catalyst.
Compose
One-card summary: SPCX on Nasdaq at $135, $85.75B deal, 5 book-runners, dual-class structure. Lockup terms and top-5 risks each in a paragraph. Ready to drop into a report section.
Recent IPOs in the last 60 days with dual-class share structures.
Retrieve
sections__ipo_screen(since="60 days ago", exclude_spac=true) → filter dual_class_flag=true.
Compose
Governance-flavored IPO cohort with voting-ratio detail. Ideagen and other governance-focused clients would treat this as a live watchlist.
Newly public companies whose use-of-proceeds includes meaningful debt repayment.
Retrieve
sections__ipo_offering for the recent-IPO cohort, filter use_of_proceeds items where purpose matches "debt" or "repay" and percent > 20%.
Compose
The "IPO to pay down PE-era leverage" cohort. Meaningful for both credit and equity research.
07 — the composed answer
Agent workflows: the whole thing.
The real power isn't any single tool call — it's the LLM composing 3-5 of them into a shaped answer.
Analyst
"Give me a competitive briefing on this newly public logistics company — whatever their filings reveal about how they compare to established peers."
Sections
ipo_screen resolves the target CIK; ipo_offering pulls their prospectus terms; ipo_risk_factors gives their classified concerns; peer_similarity against RISK FACTORS surfaces the top 5 logistics/freight companies whose risk profile they most resemble; get_filing pulls each peer's most recent MD&A.
Agent
Composes a 4-part briefing: (1) Deal structure and use-of-proceeds. (2) The three risks the new issuer flags that its peers don't emphasize. (3) The two risks its peers flag that the new issuer downplays. (4) How the peers' most recent MD&A characterizes market conditions the new issuer is entering.
PM
"Tell me what changed in every filing my portfolio put out this quarter."
Sections
For each cik in the portfolio: latest filings this quarter → filing_change against prior-year equivalent → each change bucketed by section (Risk Factors, MD&A, Legal Proceedings, Controls). Cross-referenced against the portfolio's own thesis notes.
Agent
One-page portfolio digest: the 3 filings where change magnitude exceeded 0.3 (substantive rewrite), what specifically changed, and a called-out flag on any filer that added going-concern or material-weakness language.
Researcher
"Track the emergence of climate risk in oil-and-gas 10-Ks over the last decade."
Sections
For SIC 1311 (Crude Petroleum), scan every 10-K Item 1A per year 2016-2026 → semantic hits on "climate transition risk" + "stranded assets" + "carbon regulation" → count filers per year with each theme + pull marquee examples.
Agent
A decade-long trajectory chart plus five representative disclosure texts, showing how the language evolved from "environmental regulation" (2016) to "stranded assets" (2019) to "energy transition" (2022) to "carbon pricing scenarios" (2025). Both the emergence and the framing shift are visible.
Reasonable caveats
- Coverage boundaries. Full-corpus 2020-2023 is 424B4-only for now (from the IPO-lockup extension). Full coverage across those years would need a separate ~4-6 day backfill. 2016-2019 and 2024-2026 are complete across all forms.
- Filer variance. Small-cap filers omit sub-section styling that our splitter relies on; the section-level fallback catches most of them but some residual is genuinely empty content. Named-issuer coverage (top ~2,500 filers) is near-perfect.
- Labels help, they don't replace reading. The pre-classified labels get you to the right cohort at query time. The body text is still what the analyst reads — and what the LLM should quote from in any report.
- as_of is real. Every screen result carries
provenance.as_of_effective; the corpus supports point-in-time replay for backtesting research, not just latest-state queries.