kitems  /  sections pilot report
2026-07-22 pilot complete

The kitems3 modernization pilot ran on 870 filings from June 2026 and landed with zero drift between postgres and Qdrant.

The old dragon.query grammar Ideagen relies on — proximity, stems, phrases, section-scope — passes through end-to-end on our own extraction off s3://sec4-parse-raw/. Boolean search on 758k paragraphs stays sub-second; semantic sits at ~120 ms warm. What follows is the pilot in numbers, then eight real queries.

758,155
Paragraphs indexed
postgres kitems_paragraphs · Qdrant sections_kitems_paragraphs
836
Distinct filings
10-K · 10-Q · S-1 · 424B4 · 20-F (+ amendments)
0drift
Postgres ↔ Qdrant
every row = one point, deterministic UUID5 keys
Scope

All five in-scope form families landed

The refresh timer's form set — 10-K, 10-Q, S-1, 424B4, 20-F — plus their /A amendments, S-11 variants, and a lone 10-KT transition report. 817 upserted, 52 no-sections (mostly shelf 10-Ks and empty micro-amendments), one transient Qdrant gRPC timeout that will retry next run.

Form Filings Paragraphs Avg chars Share
S-1/A123245,168288
S-1100159,747311
424B461124,013314
10-K15686,048334
20-F3362,940243
10-Q31158,864309
10-K/A239,840330
S-1124,078288
20-F/A62,888345
10-Q/A192,776275
S-11/A11,525231
10-KT1268242
Queries in the wild

The dragon grammar passes through, unchanged

Real client-shape queries against the 758k-paragraph corpus. Each card shows the query as sent, the postgres tsvector it's translated to (when different), applied filters, the top result, and end-to-end latency. Legacy wildcards (word*, word!), proximity (W/N, ADJ/N), phrases ("..."), boolean operators (AND/OR/NOT), and section-scope (/p, /s) are handled at query time.

Ideagen-style pledged-shares sweep — proximity on foreign-issuer 20-F Item 6

boolean424 ms
(pledge* W/20 share*) OR (pledge* W/20 equit*) OR (held* ADJ/10 brokerage*)
(pledge:* <20> share:*) | (pledge:* <20> equit:*) | (held:* <10> brokerage:*)
form=["20-F", "20-F/A"]
Four Seasons Education (Cayman) Inc. 20-F · Item 16K · Spousal Consent Letter score 0.0052

Pursuant to the spousal consent letter executed by the spouse of the shareholders of our VIEs, each of such spouse unconditionally and irrevocably agreed to the execution of exclus…

Natural-language semantic — going-concern language, scoped to 10-Q Part II Item 1A

semantic120 ms
substantial doubt about ability to continue operations
kitem="II-1A" mode="semantic"
Borealis Foods Inc. 10-Q · Part II · Item 1A cosine 0.874

Substantial doubt about our ability to continue as a going concern has not been alleviated.

Boolean + Porter stem — cybersecurity incidents scoped to Item 1A

boolean252 ms
cyber:* & (breach | incident | ransomware)
item="1A"
Macquarie Infrastructure Fund, L.P. 10-K · Part I · Item 1A · Business Risks rank 0.308

The Fund and its service providers, as well as the Portfolio Entities and their service providers, are susceptible to operational and information security and related risks of cybe…

Quoted phrase + AND — proxy peer-group changes in 10-K

boolean149 ms
"peer group" AND (modified | updated)
peer<->group & (modified | updated)
form=["10-K", "10-K/A"]

No paragraphs matched in this month's corpus. Query surface still resolved in 149 ms — the tsvector index short-circuited on the phrase constraint before ranking cost anything.

Recent browse — no text query, show latest 10-Q MD&A paragraphs

recent159 ms
query = None
kitem="I-2" form=["10-Q", "10-Q/A"]
APOGEE ENTERPRISES, INC. 10-Q · Part I · Item 2 · Forward-looking statements order by filing_date DESC

This Quarterly Report on Form 10-Q, including the section Management's Discussion and Analysis of Financial Condition and Results of Operations, contains certain statements that a…

Hybrid — boolean + semantic merged via reciprocal-rank fusion

hybrid339 ms
competitive pressure from artificial intelligence
form=["10-K", "S-1", "S-1/A"] mode="hybrid"
Powerfleet, Inc. 10-K · Part I · Item 1A · Macro / AI exposure rrf 0.0164

The development and deployment of AI technologies are areas of intense competition, and our competitors or other third parties may develop or deploy AI capabilities more quickly or…

Faceting

Grouping by form, kitem, and filer — all sub-300 ms

kitems_group(query, group_by, …filters…) hands back facet counts over the same filtered set search would return. Filer facets auto-attach the most-recent filer_name label so consumers don't need a second lookup.

group_by = form
Full-corpus distribution
221 ms · 758,155 paragraphs
S-1/A245,168
S-1159,747
424B4124,013
10-K86,048
20-F62,940
10-Q58,864
group_by = kitem
Where "cyber:* & breach" hits, by section
150 ms · 915 paragraphs matched
RISK FACTORS419
I-1A225
II-1A137
I-1C61
3 (20-F)48
16K (20-F)25
group_by = filer
Top mentioners of "securit:*"
247 ms · 311 paragraphs
Z Squared Inc.115
Factorial Energy52
Hyperliquid Strategies50
Hadron Energy49
Csquare, Inc.45
Perf profile

Sub-second across every mode at pilot scale

Latencies measured against the live 758,155-paragraph corpus, warm process. Semantic includes E5-base-v2 inference for the query; hybrid runs both sides + reciprocal-rank-fusion merge.

Mode Backing Representative query Latency
booleanpostgres · tsvector + GIN + ts_rank_cdcyber:* & (breach | attack | incident)559 ms
booleanpostgres · with kitem filtercyber:* & breach | attack + item=1A252 ms
booleanpostgres · legacy dragon syntaxIdeagen pledge sweep (proximity)333 ms
semanticQdrant · E5-base-v2 cosinegoing-concern natural language120 ms
semanticQdrant · with kitem filtergoing-concern + kitem=II-1A118 ms
hybridboth · reciprocal-rank fusioncompetitive AI pressure339 ms
recentpostgres · ORDER BY filing_date DESClatest 10-Q MD&A paragraphs159 ms
grouppostgres · GROUP BY (any facet)form distribution across corpus221 ms
What this proves

Parity on the pieces Ideagen and DealPoint actually use

Next

Full 2026 corpus, then MCP wrap for the agent