Strip out every insider-focus query and two-thirds of what remains is one client's accounting sweeps — but the leftover third, read by document type and buyer, exposes lanes with something the insider product's rivals lack: several independent customer segments already pulling the same filings by hand.
A structured proxy dataset — director slates & election outcomes, say-on-pay results, shareholder & activist proposals, compensation packages — extracted from DEF 14A and kitems3 sections. The single strongest new signal in the residue.
The cleanest net-new corpus: a family of enforcement/regulatory feeds nobody is productizing — coded into an action feed (type, respondent, allegation, resolution) and a comment-letter topic tracker.
A proven academic market that competes with none of our current clients.
Clause-level search & extraction over the agreements corpus — change-of-control, indemnification, MAC clauses, precedent language — for transactional and M&A practices.
Higher build effort — clause extraction is the hard part, and the natural place to prove inferno on nuanced text.
Restatements, insider pledging, changes-in-estimate, cyber (8-K 1.05), auditor changes — the dominant term clusters. Not new to our accounting reseller, but the same academics buying enforcement data want these as standalone feeds. Weigh against competing with a paying client (see the "Compete" sheet).
| Lane | Buyer breadth | Build effort | Client conflict | Verdict |
|---|---|---|---|---|
| Governance & Proxy | Medium | None | Pursue | |
| Enforcement & Regulatory | Low–Med | None | Pursue | |
| Contract / Clause Mining | Med–High | None | Watch | |
| Accounting datasets (broadened) | Medium | Ideagen | Guard |
Governance & Proxy and Enforcement & Regulatory are the picks: broadest, most diverse buyer bases, cleanly separable corpora we already index, and zero conflict with existing clients. Enforcement is the lowest-effort first build and lands in a proven academic market; Governance/Proxy has the widest commercial pull (law + exec-search + activists).
Next concrete step: point inferno + OpenRouter at one — pull the real query shapes, run an extraction pass over the matched sections, and grade precision against a hand-checked sample. One pilot proves the "search → dataset" machine, then it repeats across lanes.
Raw term volume is dominated by one client's accounting sweeps, so it understates the new lanes. Buyer breadth — how many independent segments pull a document type — is the truer demand signal here, which is why the ranking leads with it, not with volume.