MARS contact + wealth atoms inventory

Author: mars (Claude, on kee's behalf) Date: 2026-08-01 Purpose: Strategic inventory of person/company contact-adjacent atoms we already hold (or already ingest raw) vs. the latent atoms sitting in raw SEC/FEC/990/property records we haven't systematically extracted. Framing: where can MARS credibly extend into ZoomInfo/Preqin-adjacent contact intel without breaking the law + without dropping the institutional-finance positioning.


TL;DR


Current-state inventory

mars.persons_v2 — 1,889,761 rows

Field Populated % Notes
cik 211,318 11.2% SEC strong-ID
linkedin_url 45,997 2.4% Bios team ingested; partial
email 0 0% Column exists
photo_url 0 0% Column exists
primary_name ~100% Identity

Bios: 716,258 bio rows across 610,071 unique persons (32.3% of persons_v2). 20,742 long-form (>500 chars) via ADV Part 2B extraction.

mars.companies_v2 — 896,078 rows

Field Populated % Notes
primary_domain 89,705 10.0% Weak coverage vs identity
city 198,800 22.2%
country 240,336 26.8%
cik 767,096 85.6% Form D mint drove this
ticker 8,474 0.9% Atlas bridge

No address, zip, phone, or email columns on companies_v2 today.

mars.investors_v2 — 101,034 rows

Contact-adjacent fields: website, country, crd, adv_cik. Firm office address NOT extracted from ADV Part 1A into a queryable column (the raw ADV filings have it).

mars.foundation_officers_v2 — LIVE

person_name_raw + compensation_usd + other_compensation_usd + title_raw. Officer addresses NOT extracted even though 990-PF filings carry them in the raw text.

Deeds vertical — scaffolded, mostly empty

Only fec_committees_v2 exists in mars. Planned tables (per task #301): fec_contributions_v2, property_records_v2, voter_registration_v2. Zero ingest.


Latent atoms — sources we already own raw

Everything in this table is public + legally usable. The gap is extraction into queryable v2 columns, not access.

Source Substrate we own Volume What's extractable (public)
Form 4 mars.form_4_transactions_v2 + raw filings 11.2M txns / ~200K reporting persons Reporting-person address (Item 3 mandatory); relationship to issuer; officer/director flag
Form D mars.form_d_filings_v2 + raw 354K filings Issuer address (Item 3); up to 10 related-persons w/ addresses + titles
Schedule 13D/G mars.schedule_13d_v2 / _13g_v2 + raw ~16.5K filings Filer address per Item 2; group members
Form 144 mars.form_144_v2 + raw 114K filings Seller name + address (Item 1); broker; issuer relationship
ADV Part 1A adv postgres db (adrian) 52K RIA firms Firm principal-office address; branch offices; Schedule A executives w/ CRD; ownership >25% w/ addresses
ADV Part 2B s3://adv-brochures-v3 (208K PDFs) 30K persons w/ CRD Individual bio (already extracted via bios); education; disciplinary; supervisor
990-PF credit(?) or unowned raw 27M officer rows in foundation_officers_v2 Officer name + address + compensation + title (already have comp; address is NOT extracted)
990/990-EZ broader nonprofit universe ~1.5M annual filings Same shape as 990-PF
DEF 14A insider vertical ingested education extracted Exec compensation, board bios, some addresses
Press releases GDELT + Benzinga articles in news landing zone 2.4M articles Named individuals + titles + affiliations (Qwen extractable at $0.15/day scale)
Company websites mars_spider fetches ~15K RIA firm sites via bios Team pages, contact pages (spider already extracts bios; addresses often on contact pages)

Latent atoms — public sources we do NOT ingest yet

Source Volume Cost to add What's extractable
FEC individual contributions ~15-20M/yr (>$200 threshold) Free bulk download; parse effort Name + full mailing address + occupation + employer + amount + committee + date
County property records ~150M parcels US Per-county; some free, some paid; commercial wrappers exist Owner name + property address + assessed value + parcel details
Voter registration files State-by-state; some fully public Per-state; some free, some ~$500 admin fee Name + address + party + registration date + voting history
Court records (PACER + state) Federal + state Per-filing fee via PACER; state varies Litigation as HNW proxy; bankruptcy filings; divorce (some public)
USPTO patent inventors ~5M patents Free bulk Inventor name + address (weak signal, older data)
USPTO trademark filings ~2M active Free bulk Owner name + address

Latent atoms — legal enrichment from web

Source Volume feasible Cost Notes
LinkedIn URLs at scale ~100K/month via Brave+Qwen $0.02/lookup ≈ $2K/mo Public LinkedIn profile URLs; not private profile data
Corporate press releases Already in landing zone $0 marginal Extract named individuals via Qwen (already doing for lawsuit vertical)
Executive photos Public press + team pages Web-fetch effort photo_url column exists, unpopulated
Company team pages Spider partially extracts (bios) Already in pipeline Address + phone often here (contact pages)

Legal legend

Category codes for each atom class:

Nothing in this doc requires ingesting Restricted sources. Every atom listed is P / P-PII / W.


Wedge positioning

ZoomInfo Preqin/PitchBook MARS potential
Buyer Head of Sales / Marketing at B2B software co Investment analyst at fund Family office / RIA / wealth manager / prospects team at fund
Core value Contact + intent for outreach Deal + fund + LP data Contact + wealth + relationship graph for HNW targeting
Person atoms Email + phone + title + tech stack Fund + firm + role Bio + role + insider history + board seats + property + FEC + foundation
Wealth signals Employer size proxy Fund size + track record 13F AUM + insider sales + board comp + property value + foundation giving
Buyer segment B2B SaaS growth (SDR-driven) Institutional investors Wealth management + private-client outreach
Price tier $15-150K/yr per team $30-300K/yr per user $50-500K/yr per firm (RIA/family office band)

The wedge: MARS doesn't compete with ZoomInfo on "which company just raised" (Crunchbase does that shallow) or "who's the VP of Marketing at company X" (ZoomInfo owns this). MARS can uniquely own "which HNW individual just had a liquidity event, where do they live, what do they give to, where's their next investable dollar going" — because the wealth data + the identity graph + the SEC + FEC substrate already sit next to each other. Nobody has all three in one system.


Roadmap options (ranked by ROI)

Tier 1 — cheapest wins, biggest coverage bumps

These are pure extractions from raw we already ingest. Estimated effort: 1-3 days per atom, cost ~$0 (Qwen extract, no proxy needed for SEC/FEC raw).

Atom Effort Rows unlocked
Form 4 reporting-person addresses → persons_v2.address 2d ~200K unique persons
Form D issuer + related-persons addresses → companies_v2.address + persons 3d ~354K issuers + ~50K related persons
Schedule 13D/G filer addresses 1d ~16K entities
Form 144 seller addresses → persons 1d ~50K unique sellers
ADV Part 1A firm office addresses → investors_v2.address 2d ~52K RIA firms

Total tier 1: ~1 week, unlocks address on ~250K persons + ~400K companies/investors. This is the cheapest wedge into "we know where the wealthy live" territory.

Tier 2 — deeds vertical population

The scaffolded vertical needs execution. Kee owns this decision.

Atom Effort Rows unlocked
FEC contributions historical + ongoing 2-3w ~200M contributions, ~50M unique donors
Property records (top-10 states first: FL, CA, NY, TX, IL, MA, WA, CT, NJ, VA) 4-6w ~50M parcels for wealth-relevant states
Voter registration (opportunistic per state) 2w for automatable states ~100M records for cheaply-public states

Total tier 2: ~2 months of concentrated deeds work. Unlocks the "HNW targeting" wedge fully.

Tier 3 — enrichment layer

Post-tier-1 + tier-2 gaps.

Atom Cost Notes
LinkedIn URLs at scale ~$2K/mo Brave+Qwen validate
Executive photo_url ~$500/mo Web-fetch from press
990-PF officer address extraction 1w Qwen 27M officer rows
Court records / litigation Complex PACER cost + state variance

Tier 4 — cross-atom products

Once tiers 1-2 land, the atoms compose into products:

Product line naming TBD; positioning is "the wealth-intel MCP tool" not "a mini-ZoomInfo".


What this inventory is NOT

Recommendation

Start with tier 1 — Form 4 reporting-person addresses. Highest return per day of effort, uses substrate we own end-to-end, unlocks ~200K HNW-adjacent persons with addresses (SEC insiders are by definition wealth-relevant), and gives us a real dataset to validate the "wealth targeting" product hypothesis before committing to the tier 2 deeds ramp.

If Form 4 addresses validate the demand hypothesis (prospects team + Raul consume it, a customer conversation moves), then queue the rest of tier 1 as a two-week sprint, then commit to tier 2. If Form 4 doesn't find a consumer signal, the wedge is smaller than framed and we don't over-invest.

Cost through end of tier 1: ~1 week of engineer time, ~$0 in extraction compute. Cost through end of tier 2 deeds: ~2 months + FEC/property per-state costs ~$5-10K total. Compare: ZoomInfo builds this at $500M+/yr in data-acquisition costs. We start with what's free + public.