The evidence, in full

This is the nerd page. Everything the front page summarizes is derived from the record below: how the engine works, how it's measured, every number with its raw run attached, and the exact commands to reproduce all of it on your own machine. If a claim on this site can't be traced to a file in the repository, treat it as false and tell us.

v1.0 record — 118-document battery + the utility axis, measured 2026-07-19 → 2026-07-24, engine cb3627c, published as drawn

1 · The method

Two independent detection layers over one deterministic masking core. The design rule throughout: recall lives in code; the model only classifies and nominates under verbatim verification. Anything the model says that can't be found verbatim in the document is refused, not guessed at.

Layer 1 — the span pass

Layer 2 — the rails

The union, the mask, the floor

2 · The measurement protocol

A number without its denominator and its sampling story is marketing. Here is ours, stated before the results so it can't be quietly adjusted after them.

Corpus construction

What we count

What counts as a leak

Any core-identity occurrence present in the exported text. No partial credit: a surname surviving inside a compound, a glued print artifact, a shorthand in a footer — all leaks. The gate's floor is zero core leaks; one leak exits red.

3 · Results

Every number links to its raw record in the repository. Nothing is blended: in-sample and out-of-sample are separate rows, and any tier whose misses a fix set ever drew on carries its saturation label — that is what the label column is for.

Headline — the v1.0 battery (118 documents, engine cb3627c, 2026-07-21)

TierDocumentsEntitiesRecallSample status
corpus — the ratchet floor 8243/243100.00%in-sample, six generations unbroken
long-form 79k words 2260/26299.24%in-sample
walk-forward round 1 578/78100.00%rails-saturated
walk-forward round 2 5258/26099.23%mostly saturated
walk-forward round 3 36610/66791.45%partially saturated — fix set 4 drew on its misses
walk-forward round 4 441,340/1,39496.13%partially saturated
round 5 — the headline 18698/73594.97%pure out-of-sample, pre-registered publish-as-drawn
battery total 1183,487/3,63995.82%

Round 5 by occurrences: 4,018/4,093 printed occurrences masked (98.2%) · 9/11 shorthand. The walk-forward curve, final for v1.0: 84.62 → 95.38 → 88.46 → 93.76 → 94.97%. Pre-registration (FINDINGS.md, recorded 2026-07-21 before any round-5 document existed): "v1.0 publishes round 5's number WHATEVER IT IS" — honored. Junk is counted raw, not yet as a rate: 47 flagged suspects across the internal ratchet, 0–9 table rows per document on round 5 (28 total); a denominator-honest junk rate is queued and stays unclaimed until it is measured.

Per-class breakdown

Per-class rollups exist for the two exhaustively-annotated in-sample corpora. No per-class aggregate has been computed for the out-of-sample rounds yet — round-5 misses carry class tags individually (see the leak ledger below), and an out-of-sample per-class rollup is queued; until it lands, no per-class number here should be read as an out-of-sample claim.

Identity classhardening matrix — 33 docs, in-sample79k long-form — in-sample, after fixes
Persons143/14335/35
Companies & organizations190/190160/162
Brands & marks68/6824/25
IDs & accounts75/7521/22
Addresses64/6415/15
Phones49/493/3
Emails30/30

Baselines — same documents, same metric

An entity counts only when every printed variant of it has zero residual in the export — the strictest scoring direction, applied identically to every system below.

SystemScopeResult
Microsoft Presidio, as shipped 90 walk-forward documents, our ground truth 1,034/2,399 — 43.10% (per round: 51.3 · 33.5 · 33.9 · 48.9%)
Our engine, same 90 documents rounds 1–4, same metric 84.62 · 95.38 · 88.46 · 93.76%

The utility axis — round 5: does the analysis survive?

Everything above measures what the export withholds. This measures what it still carries — the number that matters when the redacted copy's whole purpose is to be analyzed by a frontier AI that was never allowed to see the original. Protocol (2026-07-24, harness committed): for each round-5 document a question-writer read the original and drafted exactly 5 analytical questions — obligations, procedure, outcomes, remedies, risk, findings; any question answerable by a name, date, address or other identity is banned, and because the writer saw only the original, questions cannot skew toward what a masked copy answers well. Two blind analysts then answered independently — one from the original, one from only the masked export, placeholders declared legitimate actors — and a judge scored each pair for substance equivalence. utility = (equivalent + ½·partial) / questions. The masked side is the round-5 export byte-for-byte as scored above — the same files that measured 94.97% / 98.2%.

MeasureResult
Aggregate — 90 questions across all 18 documents 0.983 — 87 equivalent · 3 partial · 0 divergent
Documents where every answer matched (5/5 equivalent) 15 of 18
Questions where the masked reader contradicted the original 0

Same rule as the leak ledger: every non-equivalent verdict publishes, mechanism named. govdoc-cfpb — the masked reader could not name the bankruptcy forum or docket; the forum is an identity and was masked by design — the question brushed the boundary the metric excludes. ocr-mapleton1965 — the net-to-Surplus figure and two column totals were unreadable: an over-mask (three summary dollar figures caught by an ID rail) — exactly the class of row the mandatory review exists to unmask before export. transcript-r5-3 — with the company name masked, the analyst tied a divestiture rationale to the wrong business line; industry context rides in names. Judge reasons verbatim, all 90 question/answer/verdict transcripts, and the harness as run are committed: RESULT.md · v1-round5-result.json · harness-v1.js. Analysts and judge: Claude Opus at high reasoning effort, 72 agent runs, zero errors. Honest bounds: five questions per document sample its substance, they do not exhaust it; the judge is the same model family as the analysts; identity questions are excluded by design — a redacted export deliberately cannot answer “who”, and that is the product working, not a loss. The Singapore document's answer transcripts are pruned from the committed artifact under the same redistribution ruling as its source text; its verdicts (5/5 equivalent) are retained and the recount is unaffected.

The leak ledger — round 5, every miss

Every missed entity of the headline round, verbatim as recorded in round5-results.json: 37 core + 2 shorthand. Family labels come from the round's miss taxonomy (RESULT.md); rows marked — carry no taxonomized family. Diagnosis and fixes belong to fix set 5, provable only on round 6's fresh documents. This table is the reason to trust the ones above.

Missed spanDocumentClassFamily
Amendment No. 1employment-r5-1
001-35231employment-r5-2
10.2employment-r5-2
10.3employment-r5-2
104employment-r5-2
Amendment No. 1 (as printed)merger-r5-1
AMENDMENT NO. 2 TO AGREEMENT AND PLAN OF MERGERmerger-r5-2
Bancroftsocr-goshen1961COMPANY
Science Research Associatesocr-goshen1961COMPANY
MELVIN G. HIGGINSocr-mapleton1965
Int. Harvester Co. Int. Truck Rep.ocr-mapleton1965
41022ocr-mapleton1965
2-7371ocr-mapleton1965
2-2011ocr-mapleton1965
974.102ocr-mapleton1965
the assessorsocr-mapleton1965SH
Dorothy Ballan-tyneopinion-defamation2PERSONOCR hyphen-wrapped person print
the Board’sopinion-defamation2SH
Fuseproxy-r5-1
Opening Actproxy-r5-1
S&P 500 Index (as printed)proxy-r5-1
Digital Nextproxy-r5-1
Device Care Centerproxy-r5-1
HC/SUM 474/2024sg-judgment-r5IDSG form code — rail mechanism bug, queued
HC/RA 141/2024sg-judgment-r5IDSG form code — rail mechanism bug, queued
AD/OA 17/2025sg-judgment-r5IDSG form code — rail mechanism bug, queued
CA/OA 17/2025sg-judgment-r5IDSG form code — rail mechanism bug, queued
LC carbinetranscript-r5-1BRAND
1022 Rifletranscript-r5-1BRAND
Michael Roxlandtranscript-r5-3PERSONtranscript bare-surname person
Greif Business System 2.0transcript-r5-3BRAND
Fiscal First Quarter 2025 Earnings Results Conference Calltranscript-r5-3IDevent-title ID
Austell, Georgiatranscript-r5-3ADDRESS
Fitchburg, Massachusettstranscript-r5-3ADDRESS
Alexander Waterstranscript-r5-4PERSONtranscript bare-surname person
Steven Fishertranscript-r5-4PERSONtranscript bare-surname person
Unititranscript-r5-4COMPANYbare company short-form
Gigapowertranscript-r5-4COMPANYbare company short-form
DYtranscript-r5-4IDbare ticker

Campaign log

Prior measurement campaigns, oldest first, each with its full record in the tree. History is append-only: superseded numbers stay published with what superseded them.

DateCampaignRecord
2026-07-19Founding bench — layer-by-layer measurement, chunk-geometry sweeps, the carry-list lesson EVIDENCE_2026-07-19.md
2026-07-19Hardening profile — corpus growth, stress suite, leak taxonomy, the two-engine configuration hardening-2026-07-19/REPORT.md
2026-07-20Scale ladder — two real annual reports to 79k words, before/after published unedited bench/scale/SCALE.md
2026-07-20 → 21Walk-forward programme, rounds 1–5 — the generalization record and this page's headline bench/oos-walkforward/
2026-07-21v1.0 release battery — the full 118 documents, engine cb3627c, wall-clock 12:23–21:22 bench/RELEASE_BATTERY_v1.md
2026-07-24Utility benchmark v1 — the analysis-survives axis: 90 blind original-vs-masked questions over round 5, judged for substance equivalence bench/utility/RESULT.md

4 · The environment

Every number above is attached to an exact, hash-verifiable configuration. If your hashes match, you are running what we measured.

model gemma-4-E2B_q4_0-it.gguf (stock Google QAT — no fine-tune) sha256 3646B4C147CD235A44D91DF1546D3B7D8E29B547DBE4E1F80856419AA455E6FD server llama-server --jinja --chat-template-kwargs {"enable_thinking":false} --ctx-size 8192 --parallel 1 cpu --device none --fit off --n-gpu-layers 0 (the CPU floor: no GPU required) node ≥20

The model is vanilla by construction and by check: scripts/verify-model.mjs hashes your local weights against the pin. Adaptation lives entirely in prompts-as-frozen-contracts and code rails — never a fine-tune, so the weights stay auditable against Google's published artifact.

5 · Reproduce

The corpus, ground truth, runner and floor ship in the tree. Reproduction is two commands after clone; the bench exits red on a single core leak.

git clone https://github.com/SimplerAI/simpler-redact cd simpler-redact && npm install cd app/extract && npm ci && cd ../.. # the extraction suite's pinned deps npm run verify # the gate: every measured lesson, pinned as a failing-loudly test npm run bench:all # the ratchet: full corpus through the shipped pipeline, zero-core-leak floor

The live bench needs the model running locally (see bench/BENCH.md for the recipe). Found a leak we didn't? That's a bug report, and it becomes a pinned test — the ratchet only turns one way.

6 · Claims we ban ourselves from making

Kept in the repo's constitution, enforced in review. If you catch this site making one of these, that's a bug:

  • "Provably 100% on arbitrary documents" — no such proof exists for anyone. The bench is a floor, not a proof.
  • "Set and forget" — nothing exports without human review; the review is the safety argument, not a disclaimer.
  • Blended accuracy numbers — in-sample and out-of-sample are always separate rows.
  • Retroactive denominators — the protocol above is stated before results, and the campaign log is append-only.
  • “Your AI won't notice any difference” — utility is measured and published with its partial verdicts quoted (0.983 over 90 drawn questions), never promised as lossless.