Competitive teardown · 2026-10-07

pfactorial vs. SEC Pro

pfactorial published a case study for a clause-level search engine over ~1.6M SEC & SEDAR agreements, built as a bespoke project for one information-services client. It maps almost exactly onto our agreements vertical — which already runs in production. This is where we stand, feature by feature, including where they're genuinely ahead.

pfactorial
A commissioned build for one client.

A six-phase consulting delivery (ingest → classify/NER → index → clause search/clustering → RAG → UI) for an information-services firm doing transfer-pricing, IP-valuation and regulatory research. SEC + SEDAR. Elasticsearch + a vector DB, RoBERTa/spaCy classification, OCR for scanned filings.

SEC Pro (ours)
A live platform; agreements are one vertical of it.

Clause- and exhibit-level search already serving over OpenSearch, with a classified agreements substrate in MARS, a grounded+cited legal chat, dragon proximity syntax, and plain-English query interpretation — all entity-resolved into the same graph as insiders, deals and filings.

Corpus at a glance

Their 1.6M counts SEC and SEDAR agreement filings. We're SEC-only today (SEDAR is in build) — but much deeper at the clause grain.

9.66M
SEC exhibits, full-text searchable
sec_exhibits_*
30.8M
Clause-level docs, semantic
sec_clauses_* · neural-sparse
468K
Classified agreements, 98.7% with verbatim quotes
mars.agreement_exhibits_v2
7.8M
Clause pointers with char offsets
mars.agreement_clauses_v2

Feature by feature

aheadwe lead paritycomparable gapthey lead / we're behind
Capability
pfactorial
SEC Pro (ours)
Clause-level search
Keyword matched within the relevant clause, not anywhere in the doc.
aheadSame, plus true neural-sparse semantic clause search over 30.8M clause docs — "research questions aren't keyword-shaped" is handled natively./doc-search?corpus=clauses
Keyword / boolean
Elasticsearch keyword + query expansion & boolean refinement.
aheadOpenSearch keyword + full dragon proximity (ADJ/N, w/N, /s, /p), highlighted blurbs, and plain-English mode=auto that maps phrasing to the legal term of art.corpus=exhibits · mode=auto
Grounded QA (RAG)
RAG answers grounded in the filing text.
parityShipped grounded+cited legal chat with a hard anti-fabrication contract, over the same clause/exhibit substrate via MCP.argos-chat · legal tools
Classification & taxonomy
RoBERTa contract classification + spaCy/Transformers NER + regex document typing.
parityCUAD clause taxonomy (clause_type, super_category, risk_level, favorability, survives_termination) + agreement taxonomy (agreement_type, reporting_family, agreement_kind) as query filters.#1001 · #951
Structured extraction
"Improved data extraction" — specifics not detailed.
aheadPer agreement: verbatim quotes, plain-English obligation, governing_law, parties, and 7.8M clause pointers with char offsets to deep-link the exact span.agreement_exhibits_v2 · agreement_parties_v2
Scanned / image agreements
OCR pipeline (Tesseract, PDFPlumber) — image filings searchable alongside text.
gapWe index EDGAR machine-readable text; image-only exhibits (older/scanned) are not confirmed covered. Worth verifying our coverage — their explicit OCR is a real edge.
Jurisdiction
SEC (US) + SEDAR (Canada).
gapSEC today; SEDAR in build (own workstream, not yet online). Closes once SEDAR lands.
Context & relatedness
Document clustering — related agreements grouped.
aheadEvery agreement is entity-resolved to company_id and joined to the insider / deal / filing graph — relatedness by real corporate identity, not just text similarity. (No explicit "cluster" UI — a framing difference.)
Latency
"In seconds" (qualitative only; no numbers published).
aheadMeasured: semantic clause search 1.2–1.9s, exhibit/keyword sub-second to low-ms warm. Real numbers, production SLAs.
Delivery model
One-off client build on dedicated AWS EC2.
aheadStanding multi-tenant platform; agreements reuse the same search, chat and entity graph as the rest of SEC Pro.

Where they're genuinely ahead — our to-do

Agreements are one vertical, not the product

This is the real distance. pfactorial's deliverable is an agreement search engine. Ours sits on top of everything MARS has extracted from the filings — every deal, position, obligation and person, entity-resolved to the same company_id / person_id. A clause hit joins straight to the deal it belongs to, the parties' insider trades, the debt stack, the comp. An agreement engine alone can't answer "show me the merger agreements for deals where an insider was also selling under a 10b5-1 plan" — ours can, because the agreement is already wired to the rest.

M&A deals + filing trail
Deals with their full SEC filing trail (S-4, proxy, 425, tender docs).
merger_deals_v2 · /deal-filings
Funding & pre-IPO
Priced rounds, amounts, IPO lock-up tiers.
funding_deals_v2 · ipo_lockup_tiers_v2
Debt & bondholders
Reported total debt + the itemized instrument stack.
tranche_debt_totals_v2
Insider trades + 10b5-1 plans
11.5M+ Form 3/4/5 rows; Item 408(a) plan adoptions.
transactions · rule_10b51_v2
Institutional holdings
13F positions, 1998→now (in build).
13F → MARS
Signals events (~4M)
Appointments, litigation, dividends, bankruptcies, breaches.
*_events_v2
People graph
Bios, work history, education, board & 990 roles.
persons_v2 · person_*_v2
Compensation
Named-executive & director comp per filing.
gpt_compensation_summary
Fundamentals & earnings
XBRL financials; earnings calendar, surprise & drift.
fundamentals_cooked · earnings_*_v2

Plus curated news (~30M articles), ADV advisers, Form C/D, and more — the same extraction layer, one entity graph. pfactorial's case study scopes to agreements alone.

Bottom line

pfactorial hand-built, for a single client, roughly the capability we already run as a platform: clause-level semantic + keyword search over SEC agreements with grounded QA. On the things that are hard to retrofit — a 30.8M-doc semantic clause index, verbatim-quote + obligation + governing-law extraction, CUAD taxonomy, dragon proximity, a shipped cited chat — we're ahead or at parity. And the decisive difference isn't the search at all: their deliverable stops at agreements, while ours is one vertical of a whole extracted-data graph — deals, debt, holdings, people, comp, fundamentals — joined to every agreement by the same entity resolution.

They have two things we don't: OCR of scanned agreements and SEDAR. SEDAR is already in build. OCR is the one worth a real look — it's the clearest differentiator in their write-up and the one place a prospect could find us genuinely short.

Source: pfactorial case study — "A purpose-built search engine for 1.6 million SEC/SEDAR agreements."
Our figures from live corpora (sec_exhibits_*, sec_clauses_* on OpenSearch; mars.agreement_exhibits_v2 / agreement_clauses_v2), 2026-10-07.