Audit Analytics doesn't buy search — they buy raw material for a dataset factory. Every high-frequency query maps to a product they sell: restatements, out-of-period adjustments, insider pledging, auditor changes. Their vocabulary is their catalog.
Audit Analytics and Ideagen are the same client — Audit Analytics was acquired by Ideagen and accounts migrated from @auditanalytics.com to @ideagen.com (verified: Kristyn Plante holds both). The heavy, current usage sits under the Ideagen domain — a large offshore analyst team in Malaysia plus the original US analysts.
| Dataset line | Their signal terms (verbatim) | Hits | Intensity |
|---|---|---|---|
| Restatements & out-of-period Non-reliance, Big-R / little-r |
restatement · out period · immaterial · misstatement · misclassification · overstate · prior period · correction | ~19,500 | Flagship |
| Changes in estimate | estimate change · depreciation · useful life · accrual | ~10,700 | Core |
| Insider pledging | pledge · margin · collateral · brokerage · equity w/20 shares | ~4,550 | Core |
| Cyber incidents SEC 2023 rules · 8-K 1.05 |
cyber · attack | ~1,970 | Ramping |
| Auditor & opinion changes | auditor · independent accountant · opinion | ~2,000 | Ramping |
Counts = occurrences of the term across their 37,768 text queries. The restatement sweep is re-run near-verbatim ~19,500 times across form/date slices.
The workflow is a discovery→hand-coding funnel: analysts run a small canon of standing boolean sweeps (heavy on W/20/ADJ/10 proximity), pull the candidate filings, then read and code them into structured datasets. They brute-force full-text and narrow by hand rather than using section-scoped precision. Form breadth is unusually wide — 10-K, 20-F/40-F (foreign issuers), the N-CSR/NSAR fund family, DEF 14A, 8-K.
They re-run the restatement sweep ~19,500 times by hand. That is a scheduled monitor begging to exist. Convert their standing queries into daily push alerts — we already own the alerter engine and it speaks the same query grammar, so their sweeps port as-is. Recurring revenue for us, a labor line erased for them.
Their entire cost base is analysts reading matches to code fields — restatement type, period, dollar amount, reason. We already do LLM extraction elsewhere (the gpt_* pipelines). Offer auto-extracted structured output on their matched sections; they buy enrichment instead of paying to read. Highest-margin upsell, and it moves us up the value chain into their moat.
They brute-force sec4 and narrow by hand. Section-scoped kitems3 search (restatement language in the right item) cuts analyst false-positive load. This is exactly what the sections team's item-3 rebuild improves — it has a named paying beneficiary. Sell it as a precision tier; prioritize 20-F/40-F/N-CSR coverage they can't get elsewhere.
70% of their calls are metadata browse and the workflow is a factory — they want cohorts, not a search box. A firehose/API tier with higher limits and bulk export monetizes the offshore team's throughput far better than counting interactive seats.
The cyber cluster shows them standing up a product on SEC's 2023 rules. Sell a purpose-built 8-K Item 1.05 cyber feed — and the same for CAMs, going-concern, ICFR weakness. You already know they'll buy, because they're already searching for it.