smart_money · internal briefing for raul
Prep for the Q4 calls — Wed with Steve (head of data), Thu/Fri with Venkat (VP, data). Same material either way: Steve will probe the mechanism, Venkat will probe the roadmap and the honesty of the gaps. Lead with the proof, not the pitch.
The one-liner
Why this is the whole conversation
"Anyone can load 13F data — it's keeping track of all the changes that is the value." — relayed from the client, via Raul
Read that as a direct spec, not a vague preference. It splits into two things a feed has to get right:
1. Quarter-over-quarter flow — did a manager buy, sell, or hold. We've done this well for a while: share-based flow (not price-inclusive), bond-mistagging filtered out, CIK re-registrations collapsed so a manager's book doesn't fake a mass exit when they just re-registered.
2. Security identity through a corporate action — did the security itself change. This is the part that was missing until this month, and it's the part that actually differentiates from "anyone can load 13F data."
What's actually built, right now
Three pieces, stacked. Every 13F holding gets a validated identifier, cross-checked against the SEC's own record, and resolved through any known corporate action before it ever reaches a flow calculation.
Layer 1 — validation. Every CUSIP on every holding is normalized and check-digit verified (the standard ANSI X9.6 algorithm — public, documented arithmetic, no data license required) and cross-checked against the SEC's own Official List of Section 13(f) Securities, ingested for its full published history.
Layer 2 — the succession ledger. 260 real corporate actions — mergers, reverse splits, reincorporations, renames, one bankruptcy — each individually verified against a primary source (an SEC 8-K, a Nasdaq corporate-action notice, a press release) before being added. Not inferred from a pattern and trusted blindly: every row has a citation.
Layer 3 — it's wired into the live numbers, not a side table. The succession ledger sits upstream of every quarter-over-quarter calculation. A security that changed CUSIPs three separate times over five years still resolves to one continuous position, share counts correctly adjusted for every split along the way.
The output is inspectable, not a black box. A single exportable table — one row per security per quarter — shows the whole chain: reported CUSIP → normalized → checked against the SEC list → resolved to today's identifier → the specific event that caused any change, cited. Someone at Q4 can pull it and check our work directly.
Two concrete cases, if they want to see it work
1-for-5 reverse split, Jan 2026. Share counts are re-expressed at the correct 5:1 ratio automatically — the holder never actually sold anything. Independently confirmed: the SEC's own official list marks Amcor's old CUSIP 'DELETED' the following quarter. We didn't have to take our own word for it.
Real, continuous trading activity — a decrease then an increase — survives the merger intact instead of vanishing into an exit/new pair. The same mechanism composes across multi-hop chains: one security in our data changed CUSIP three separate times (a merger, then two later splits) and still resolves to a single position at the correctly compounded ratio.
What we're honest about
Data people trust a vendor who names their own gaps more than one who claims none. These are real, current, and each has a next step.
The detector surfaces candidates automatically from filing patterns; a human (or an independently-verified research pass) confirms each one before it goes live. 260 is the high-confidence tier, cleared first. The rest is a scale problem, not a correctness problem — same process, bigger pool.
SPAC units splitting into common stock + warrants, and preferred-to-common conversions. Both are genuine "did the position's nature actually change" questions where the honest answer is "it depends," and we chose not to guess. Held out until there's a real policy call, not fabricated.
We verified our cross-check against the list is working correctly (Microsoft matches every single quarter back to 2019, zero misses). But the list itself doesn't cover every legitimate security every quarter — some ADRs, foreign primary listings, thin small-caps. We treat absence from the list as a soft signal, not proof of a problem, and we say so in the data itself.
That's the actual moat FactSet and others pay for, and it's why they can promise near-total coverage. We're reconstructing the event graph from free public sources — SEC filings, exchange notices, press releases. Real recall, honestly bounded, strongest for anything with any public paper trail.
Anticipated questions
Every one of the 260 events has an independent, citable source — an SEC 8-K filing, a Nasdaq corporate-action notice, or a company press release — checked against the pattern, not inferred from it. We caught and rejected candidates during curation where the pattern matched but the underlying event turned out to be something else (a SPAC unit splitting, a preferred-to-common conversion) — those got excluded rather than force-fit.
Ongoing. The candidate-detection step runs continuously off the same signal (paired mass-exit/mass-new activity on the same issuer in one quarter); curation is the human-verification step we run in batches. The mechanism itself doesn't need touching — new events just get added to the same ledger.
Yes — that's exactly what the audit-trail export is for. Reported CUSIP, normalized form, check-digit result, SEC-list match, resolved identity, and the specific cited event, one row per security per quarter. Nothing about the resolution is hidden inside a black-box join.
Holdings data covers 2020 forward as the reliable core (with some earlier data present but not the primary scope). The SEC's own official securities list — our cross-check layer — covers its full published history, 2019 Q1 through today.
We don't have a clean global number — that's precisely the problem naive feeds have, since a merger or split doesn't announce itself as an error, it just silently produces two fake events instead of one real one. What we can show is the before/after on real, verified cases: the fake-exit-plus-fake-new pattern disappears entirely for every one of the 260 curated events, and a residual detector confirms it — the false-signal rate on those specific securities dropped from thousands of flagged quarters to about 1%.
Evaluate it directly — the audit-trail export is built for exactly that. Give us a handful of tickers you already know the corporate-action history for, and we'll show you what our data says versus what happened.
If you only say one thing
Don't oversell the backlog or the edge cases — data people respect the caveats more than the pitch. The mechanism is real and it's already live in production, not a demo built for this call.