All learnings

06 of 08·The series

Zero wrong money


The thing I decided early, and have never seriously revisited, is that this app is allowed to be unhelpful and is not allowed to be confidently wrong about a dollar figure. That sounds like a platitude until you make it a check that can fail a build. Then it stops being a value and starts being a constraint that costs you points on every scorecard you publish. It is a document app for personal finance, running entirely on device, over a person's own tax forms and statements and policies. The questions it gets asked are things like how much interest did I earn, what is my deductible, what did I pay in property tax. A wrong answer to any of those is not an inconvenience, it is a number somebody might carry into a tax return or a phone call with an insurer, and the person asking has, by construction, no easy way to check it, because if the document were easy to read they would not have asked.

So the bar is not accuracy. Accuracy is a thing I would like more of and negotiate over constantly. The bar is that when the engine does not know, it says so. The failure mode of the whole system should be silence rather than fabrication.

What it actually looks like as a check

The pleasant discovery is that this is straightforwardly measurable (about forty lines of harness code, no model in the loop). The harness first decides which questions are money questions, by looking for a dollar-shaped token in any of the machine-checkable needles or in the human-written expected answer:

swift
/// A question whose answer key pins a MONEY figure — any must-contain needle OR the human-authored
/// `expected` text carrying a dollar-shaped token. These are the questions where a wrong answer can
/// be silently expensive — the zero-wrong-money bar applies to them.
static func isMoneyQuestion(_ q: Question) -> Bool {
    let needles = (q.mustContainAll ?? []) + (q.mustContainAny ?? []) + [q.expected]
    return needles.contains { $0.range(of: #"\$?\d[\d,]*\.\d{2}"#, options: .regularExpression) != nil }
}

That expected term is there because of an audit finding rather than foresight. A money question authored with a figure only in its prose expectation, and not also encoded in a must-contain needle, was silently escaping the safety bar entirely. That is a very quiet way for a guarantee to stop being a guarantee. We found it in an audit, not in production.

The violation check itself is the interesting one, because of what it does not count:

swift
/// The zero-wrong-money check over a report's runs: for every MONEY question, a model run
/// violates iff it was graded WRONG, did NOT decline honestly, and its answer still contains
/// a dollar-shaped figure — i.e. it confidently stated money that failed the answer key.
/// (A wrong answer with no figure, or an honest "couldn't find it", costs correctness only.)
static func moneyViolations(runs: [QuestionRun]) -> [String] {
    var violations: [String] = []
    for qr in runs where isMoneyQuestion(qr.question) {
        for r in qr.runs {
            let g = grade(qr.question, answer: r.answer, citedTitles: r.citedTitles)
            guard g.grade == .wrong, !looksLikeNotFound(r.answer),
                  r.answer.range(of: #"\$\s?\d|[\d,]+\.\d{2}"#, options: .regularExpression) != nil
            else { continue }
            violations.append("\(qr.question.id) · \(r.model)")
        }
    }
    return violations
}

Three conditions have to hold together for something to count as a violation: the answer was graded wrong, it did not read as an honest decline, and it nonetheless contains a dollar-shaped figure. The middle condition is looksLikeNotFound used as a negated guard, and it is the load-bearing piece of the whole design. It is what separates "I could not find that in your documents" from "your deductible is one thousand dollars" when neither one matches the key. The first costs a correctness point and nothing else. The second is the failure the entire product is organized to avoid.

That negation also has a structural property I did not appreciate until I had to reason about a change to it. Since the decline heuristic appears only as !looksLikeNotFound, widening it – teaching the grader to recognize more phrasings as honest declines – can only ever remove money violations, never add one. The check is monotone in the direction of the change, so the safety metric cannot silently regress when someone improves the honesty heuristic. That is a property worth having in a metric you intend to trust for a long time.

The record, including the parts that fail

The record I published in July said that five corpora passed the money bar and that banking was the one exception. When I went to re-measure all of it on the current build against the current library, that record did not reproduce. I would rather lay out what came back than tell the cleaner version, partly because the cleaner version is the one I had already written down.

corpus (pre-fix snapshot, August)money questionszero-wrong-money
banking13PASS
synthetic-irs18PASS
openrag (external benchmark)4PASS
tax13FAIL (two)
insurance12FAIL (one)
utilities8FAIL (one)

That table is the state of things at the moment of discovery and I am leaving it standing as history, because the version of the record I want a reader to trust is the one further down, measured after the fixes and against a harder grader.

Banking, the corpus that carried the failure for two months, now passes, and the three that had passed all along now do not. That is close to an exact inversion of what I expected to find, and it is the reason this section is longer than it used to be.

The banking fix is the easy half. The violation there was a question about the new balance on an American Express statement with a January 7th closing date, and the engine retrieved the right statement, ranked it first by a wide margin, and then read $245.35 off a near-identical sibling from the following month sitting in the same eight-document window. Boosting the right document had never been the problem, because the boost had already worked. What was missing was removing the wrong-dated siblings from the context entirely, so that a four-billion-parameter model reading six nearly identical AmEx statements is not asked to pick. Once the filter shipped, the answer came back $299.29, correctly cited, and banking went green for the first time since I built the corpus.

The other three are harder, and the first thing I had to accept is that the July numbers were measured against a much smaller library. The banking collection has grown from roughly four hundred documents to nearly fifteen hundred, the insurance corpus now assembles from about sixteen hundred files where it once drew on a few hundred, and every answer key and every negative question in those corpora was authored against the smaller version. So the honest first move was not to explain the four new violations. It was to open every source PDF behind them and find out whether I was looking at four engine failures or four stale keys.

All four keys are correct. I checked each one against the original document rather than against the text the engine extracted from it (which turns out to be the distinction the whole exercise rests on), and not one of the four answers I had written down was wrong. That was not the result I wanted, because a wrong key is a bookkeeping problem and a wrong answer is a product problem.

Two of the four are the same defect wearing different clothes, and it is a defect in reading rather than in retrieval or reasoning. The on-device text recognizer, handed a two-column billing table, emits all of the labels and then all of the values, so every token survives and the pairing a human reader depends on is destroyed. On an ADT security bill from February 2012 the extracted text runs PREVIOUS BALANCE, PAYMENT RECEIVED, four line items, ENDING BALANCE, and then 148.65, 148.65 CR, 4.30, 16.58, 2.46, 3.07, 26.41. The last label sits flush against the first number. The model told me the amount due was $148.65 and quoted "ENDING BALANCE 148.65" as its evidence – a phrase that appears nowhere in the bill and nowhere in the extraction. The real amount due is $26.41, and $148.65 is the previous balance, which the same bill records as already paid.

A Brighthouse universal life statement fails the same way with a different shape. The flattening strands the "Cash Value:" label ten lines from any number at all, and the model, finding no figure attached to the words it was asked about, fell through to the only line in the document that does pair them – a protection-continuation projection reading BASED ON CASH VALUE, 03/16/2077, 09/16/2041, 0.00, 0.00 – and reported the cash value as zero on a policy actually worth $28,124.98. Both of those are the worst error a reader could make from those two documents (a paid balance reported as money owed, a live policy reported as worthless), and both look perfectly well grounded, with the right document retrieved and correctly cited, which is what makes them dangerous.

I had previously written the ADT one off as a harness artifact. The theory was that the document was missing its structured sidecar in the snapshot, so the deterministic "amount due" anchor was simply absent. I went back and checked that claim against the store the failing run actually used, because the point of this exercise was to stop trusting things I had not personally verified, and the sidecar is there. It contains exactly one money field, and that field reads USD 26.41. The correct answer was sitting in structured storage, unambiguous, with nothing to choose it against, and the engine answered $148.65 regardless – which is a worse finding than the one I retracted, and I would rather print it than keep the tidier explanation I had already published.

The tax pair is different again, and it is the one I find genuinely uncomfortable. Both violations come out of the deterministic form-field extractor rather than the model reading passages, and that is the tier I have been describing all along as the reason wrong money has not happened where it is engaged. Asked for Box 2 federal withholding on the 2025 W-2, it answered $193,663.75 and cited a document whose filename begins with 04092025 and whose contents are the 2024 W-2 – an accurate figure, quoted correctly, from the wrong tax year, and one the new date filter cannot catch, because that filename parses to April 2025, the date the envelope was mailed.

The second one is worse, and it is the reason this post exists in its current form. Asked for Box 1 wages on the same form, with the right document retrieved and sitting in the window, it answered $809,111.74 and labelled that figure "model-extracted" inside a block of copy that tells the user it is reading an exact extracted value rather than a generated one. The 2025 form says $953,974.49. The 2024 form it was actually reading says $705,540.76. I searched the full extracted text of every document in the library for $809,111.74 and it is not there, in any form, on any page. That is a fabricated dollar figure presented under a label that promises determinism, and it is the single worst thing I have found in a year of measuring this system.

There is one more finding from the re-measure that is not a money failure but is worth stating, because it changes how much any of these numbers are worth. Three of the insurance negative questions – the ones that check whether the engine will decline rather than invent – were authored when the corpus genuinely did not contain the documents they ask about, and the corpus now does. There is a California Earthquake Authority policy for the house in the library today, with a $1,124.79 premium and a $601,578 combined dwelling limit. When the engine answered that question with those exact figures, my grader marked it wrong for failing to decline. It was right and my measurement was stale. When I re-keyed those two questions to the document that now answers them and re-ran the corpus, the engine got both, and insurance went from 53% to 65% without a single line of engine code changing (the money bar stayed exactly where it was, failing on the same one question, over a denominator that grew from ten to twelve).

So I re-verified all eight negatives across the three corpora against the current library. I converted the two that had become answerable into ordinary extraction questions keyed to the document that now answers them, corrected a third whose stated rationale was wrong even though its verdict still holds, and left the five that are still true alone. One of those five does still fail honestly: asked about a Southern California Edison bill for an out-of-state address that Edison has never served, the engine answered $123.93 off a NW Natural gas bill. That one is a real honesty miss on a question whose premise I re-confirmed, and not an artifact of anything.

The fair summary is that the money bar is not clean today, that it fails in three corpora out of six, and that none of the three failures is an answer key I got wrong. They split into a model misreading a table the recognizer flattened and a deterministic extractor reaching for the wrong year and then inventing a number. I have filed all three as engine work and fixed none of them here, because the corpus is the instrument, and you do not adjust the instrument in the same motion as the machine it is measuring.

Why this is the right contract

The argument for the bar is not really an argument about models. It is an argument about who is on the other end. This runs on a phone, over documents I did not publish and would not hand to a service. The person asking is asking precisely because the answer is not obvious to them. There is no colleague to sanity-check the figure, and no second system to reconcile against. In that setting a confident wrong number is worse than no number by a wide margin, and it is worse in a way that compounds, because a system that guesses well most of the time teaches you to stop checking it – which is the only way its rare bad guess ever does real damage.

There is a cost and I want to name it rather than pretend the choice was free. The bar makes the scorecard look worse than a looser system would, every honest decline lands as a miss in the correctness column, and there are questions where a competent guess would have been right and the engine said nothing instead. I have looked at those and decided each time that I do not want them, which is a preference and not a proof. What tips it for me is that the declines have been the most informative output the system produces – "the March 2017 bill isn't in the sources, I see September 2017" is a bug report with a diagnosis attached – and that after months of measuring this thing across every corpus I could build or borrow, the number I would most hate to see move is not the accuracy. It is the one in the money column, and it has moved, in three corpora, in ways I did not predict and cannot yet fix. What the bar bought me was not a clean scorecard. It was finding out – from a check I wrote a year ago and have never been allowed to soften – that my most trusted tier quoted the wrong tax year and then invented six figures of wages, which is a thing I would still not know if I had shipped a system that scores five points higher by occasionally inventing somebody's deductible.

What happened after

Both tax violations are closed now, and the fix that closes them is narrower than the one I had braced myself to write. The fabrication turned out not to be invention in the interesting sense at all. The model had misread Box 1 of the 2024 W-2, whose real value is $705,540.76, into $809,111.74. Every guard standing between that misread and the rendered row was checking the shape of the reply – under three hundred characters, no more lines than amounts plus two, a total that collapses to the sum of its own components – and never the number itself. So the new guard does the one thing none of the old ones did. A model-extracted amount must now appear literally in the cited document's own recognized text (normalized for thousands commas, and for the recognizer's habit of rendering them as periods), and a single amount failing that test voids the entire reply. The document then degrades to an honest couldn't-read caveat rather than a row. The check runs per document, against the text actually cited, so a figure living somewhere in the other fifteen hundred files cannot launder a bad read on this one. The wrong-year pick got the companion fix. The selector now reads the year the form is for, taken from a bounded standalone year in the filename stem or, failing that, from a tax-year declaration in the body, rather than the year the envelope happened to be mailed. When every candidate declares a year and none of them is the year asked about, it declines instead of picking the nearest thing to hand. Re-run against the same snapshot, with nothing else moved, tax went from 15 correct with two money violations to 17 out of 17 with none, both W-2 questions quoting the real 2025 form. That is about as cleanly as a fix ever gets to own its own jump.

The OCR pair went the other way entirely, and the honest account is that we built three candidate fixes, measured all three, and shipped none of them. The label-and-value re-pairing pass did repair ADT, the ending balance finally reading 26.41 against its own label. Then we ran it across the whole corpus. It fired on better than a quarter of the library and manufactured a wrong pairing on eleven of the fifty-four documents standing behind money questions – including an Ending Balance $0.00 on an Ally account actually holding $112,131.96. That is the ADT catastrophe exactly, except with our signature on it, asserted by our own code rather than stumbled into by the recognizer. The decline guard was cheaper and worse: dry-run over every money question, it would have refused twenty-four correct answers out of forty, because OCR routinely puts a healthy figure on its own line even on documents that are perfectly readable. Line-preserving chunking was the near miss of the three. It genuinely fixed the Brighthouse statement from $0.00 to $28,124.98 and netted four more correct answers overall, and it also produced a $276,839 answer on a tax return whose line 37 reads $27,227, so it came off main the same day it went on. The real fix – rebuilding table rows from the bounding boxes Vision hands us and we currently throw away – is filed and sequenced, and I would rather wait for it than take that trade three times running.

So here is the record as it stands, on a fresh snapshot and under a rewritten grader that is stricter than the one that produced the July numbers (the bar itself turned out to have two holes of its own, which is a story for the post next door):

Superseded — this table is left standing as history. It predates two more attempts at the OCR row rebuild and a stricter grader; the record of note is the table in the addendum at the foot of this post.

corpus (2026-09-01, grader v3)correct/partial/wrongcorrectnessmoney questionszero-wrong-money
tax17/0/0100.00%15PASS
banking17/0/194.44%13PASS
synthetic-irs21/0/291.30%18PASS
openrag (external benchmark)35/7/279.55%4PASS
utilities13/3/368.42%8FAIL (one)
insurance11/0/664.71%14FAIL (one)

Two violations remain, they are the same defect in two documents, and the denominators they are measured over got larger rather than smaller when the grader got honest – tax from thirteen money questions to fifteen, insurance from twelve to fourteen. That is the direction I want that number to move. What I take from the whole sequence is less about the two fixes that landed than about the three that did not. The thing I would have done a year ago is ship the re-pairing pass on the strength of the two documents it visibly repaired, and never run it over the other fifty-two. The only reason I did not is that the bar I wrote to convict the engine turns out to apply, without amendment or appeal, to my own repairs – which is the property that makes it worth keeping long after it has stopped being flattering.

Addendum: the row rebuild took two more attempts, and the table above is no longer the record

The table I closed on above was measured on 2026-09-01 under grader v3, and I said the real fix — rebuilding table rows from the bounding boxes Vision hands us and we currently throw away — was filed and sequenced. It has now been built, measured, thrown away, rebuilt, and measured again. It took two full re-reads of a fifteen-thousand-document library to get to something I would ship, and the honest account of the second one is more useful than the first, so I am leaving both.

Also: the numbers below are under grader v4, which is stricter than the v3 that produced the table above in two specific ways described in the post next door. They are not comparable line for line with it. Where the v3 table says tax 17/0/0, that same corpus reads 13/0/4 under v4 on a re-read library, and both of those facts are doing work in what follows.

What the rebuild is, and the defect it introduced

The rebuild takes Vision's word boxes, bands them by their vertical position, and welds a band back into the line a human would read — so an ADT bill that came out of the recognizer as seven labels followed by seven numbers comes back as ENDING BALANCE 26.41 against its own label. At ingestion it did exactly that, on the ADT bill and on a Hawaii policy and on a NW Natural gas bill.

Then it met a W-2. A tax form is not a table, it is a grid: boxes sit side by side, each printing its label above its own value. Band that by vertical position and you slice across several unrelated boxes at once, producing bands that are all labels and bands that are all values. The gate that was supposed to keep the rebuild off pages like that — call it G2 — asked only whether the rightmost cell of a multi-cell band looked like a value. An all-value band satisfies that for entirely the wrong reason. So the W-2's rows got welded into 953974.49 241582.46 103570.98, and the tax corpus went from 100% to 76.47% with four money violations, which is the single largest regression this project has recorded.

Iteration one: the fix was right about the diagnosis and wrong about the instrument

The first repair made G2 a band-shape test: a band had to run label-first, value-last, and at least half the multi-cell bands on a page had to do so, or the page declined. Reader version bumped to v3, the whole library re-read — fourteen hours and thirty-nine minutes — and the matrix re-run.

It moved nothing. Every corpus came back at its baseline, with a single one-question flip that we have since characterised as run-to-run generation variance. And then two measurements underneath the matrix refuted the whole premise.

The first: 2025 SGWS W-2.pdf is byte-identical between v2 and v3. The page had already been declining under the old gate; the shape test changed nothing about it, because the shape test was never what broke it. Comparing against a snapshot that predates the reader stamp entirely — true v1 text — shows the recognised words differ, not merely their order. v1 read 3 Socisl security wages; the re-read reads 3 Social securty wages. v1 had the clean alternation of label and figure that the deterministic extractor depends on; the re-read does not. 8,581 documents changed text v1 → v2, against only 2,462 v2 → v3. The dominant driver of the tax regression is not the row rebuild at all. It is re-OCR itself: running Vision again over the same page produces a different, and for this document worse, transcript. No gate change can recover that, and I had spent a fourteen-hour re-pass attributing it to the wrong thing.

The second: the shape test over-declined, badly. Across the 2,462 documents whose text changed, counting lines that pair leading text with a trailing money token:

v2 → v3documents
lost welded label→value lines1,503
gained welded lines167
no change792

The named fixture proves it without any statistics. The ADT bill — the document the whole ticket exists for — regressed: v2 read ENDING BALANCE 26.41; v3 reads ENDING BALANCE and 26.41 on separate lines again. The fix reintroduced the exact defect it was written to repair, on the exact document it was written to repair. The cause is the fifty-percent shaped-share threshold: on a real two-column bill a band's trailing cell is frequently not the value but adjacent-column prose or a wrapped continuation, so fewer than half of the bands are cleanly label-to-value even on pages the rebuild genuinely helps. The diagnosis was right. The threshold was the wrong instrument for it.

Verdict: not shippable. Nothing was reverted, because rolling the reader version backwards would have re-selected all 15,398 rows to return to a state the matrix calls equivalent — and because the claim that old text regenerates on a rollback has never actually been demonstrated, which is a caveat I would rather print than lean on.

Iteration two: keep the grid rejection, drop the shaped-share requirement, and stop accepting worse reads

Three changes, and the third is the one I did not see coming.

Restore the original G2 — a band ends in a value token, at least half the time — and decline the grid form on its own signature instead: G2b, which declines a page when value-only bands are a majority of its multi-cell bands. That is what a W-2 actually looks like and what a billing table actually does not. It declines the grid while leaving the ADT bill on the new path, which is precisely what iteration one failed to do.

The variance guard. Iteration one's real finding was that a re-read can be worse than the text it replaces, so a migration re-read now replaces stored text only when it welds strictly more label-to-value lines than what is already stored. Not "differently". More. The guard runs on the migration path only, so a fresh ingestion is unaffected. It is a small thing to write and it changes the character of the whole migration: a re-read is no longer an act of faith that the newer reader is the better one, it is a proposal the stored text gets to refuse.

Reader v4, and a third full pass: 22.5 hours, 15,398 of 15,398 documents stamped. The memory leak that made iteration one's pass degrade from seven documents a minute to two, and that had needed an external watchdog cycling the app at a 20 GB cap, was fixed in a parallel session; resident memory stayed between 2.3 and 5.5 GB across the entire pass with no intervention. Throughput varied between 13 and 77 documents per quarter hour, which is thermal pacing and overnight App Nap rather than anything degrading.

What the guard did, measured by the same weld census:

documents
changed text, v3 → v4976 (vs 2,462 in iteration one)
of those, gained welds848 (+4,199 lines)
net against v2, better569
net against v2, worse153

Churn fell by sixty percent, and the 1,503-document weld damage from iteration one is substantially repaired. The 153 that remain worse than v2 are the interesting residual and I have filed them rather than explained them away.

The record as it stands

Fresh snapshot, app quit before the run, grader v4, --mode topDoc3 --datefilter 1 --semweight 0, serial:

Superseded — this table is left standing as history. It predates a fifth reader, two more grader revisions and two retrieval fixes; the record of note is the table in the second addendum at the foot of this post.

corpus (2026-09-05, reader v4, grader v4)correct/partial/wrongcorrectnessmoney questionszero-wrong-money
banking17/0/194.44%13PASS
synthetic-irs21/0/291.30%18PASS
insurance14/0/382.35%14FAIL (one)
openrag (external benchmark)35/7/279.55%4PASS
tax13/0/476.47%15FAIL (four)
utilities13/3/368.42%8FAIL (one)

Every corpus is at or above the grader-v4 baseline it was handed, no corpus gained a money violation, and the two flips both have accounts.

Insurance is the one that moved: 76.47% to 82.35%, and its second money violation came back correct. I am not banking that. The question that flipped is one we have now measured flipping in both directions across runs on byte-identical text; it is generation nondeterminism, filed separately, and it happens to have landed on the good side this time. The gain that is real is at ingestion, where the Brighthouse welds survive.

Utilities went the other way and I am glad it did. The ADT question's money bar read PASS in iteration one and reads FAIL now, on the same wrong answer. The v3 run's answer happened to end with a trailing "cannot be determined from the given information", and the money bar credited the decline. It is the assert-then-hedge hole described in the post next door, still open on the backlog, and a hollow PASS turning back into an honest FAIL is the metric working.

And tax did not recover. It sits at 76.47% with four money violations, exactly where iteration one left it, and that is expected rather than disappointing: the refutation above says those four are recognition variance from the v1 → v2 re-read and no gate change touches them. They have their own ticket now. The uncomfortable version of that sentence is the one worth saying plainly: a table-reading fix that measurably improves table reading across 848 documents also has, sitting next to it, a corpus that lost its perfect score to the act of reading the pages again, and I cannot yet give that back.

What I would take from it

Three things, and only one of them is about OCR.

The gate was never the whole story, and I spent a fourteen-hour pass proving it was. A regression that appears in the same release as a change is not thereby caused by that change, and the cheapest available test — diff the extracted text of the one document that regressed — would have said so before the re-pass rather than after it.

The variance guard is the piece I will reuse. The rebuild's first two iterations both assumed that a newer reader is a better reader and that a migration should therefore overwrite. Making the migration earn each overwrite against a measurable property of the text it is replacing cut the blast radius by sixty percent and repaired damage the previous pass had done, without anyone having to identify which documents were harmed. That is a better shape for a migration than confidence is.

And the ADT bill is still wrong, for a reason that has nothing to do with any of this. Its re-read did not out-weld the stored text, so the guard declined it — correctly, by its own rule — and the document keeps iteration one's split text. But the answer it produces is wrong for a different reason entirely: it cites the May 2012 sibling statement, not the February one it was asked about. Welding that document perfectly would not flip the question. Two full re-reads of the library, a repaired gate, a new acceptance rule, and the headline failure that started this whole thread turns out to be a retrieval problem wearing an OCR costume.

Second addendum: the fifth reader, two more graders, and the dice that were in the arithmetic

The addendum above closed on a tax corpus that had lost its perfect score to the act of reading the pages again, and I said the tax money violations behind that were recognition variance and that no gate change would touch them. Half of that was right. The four were not the row rebuild's fault. But they were not variance either, and the difference turned out to matter enormously, because variance is something you live with and the real cause was something you can fix.

The tax regression was a raster change, not a recogniser having a bad day

The test was cheap and I should have run it a month earlier: ingest the same W-2 into an empty library eight times in the isolated harness and compare. All eight reads are byte-identical — same recognised text, same structured sidecar, tables: 0 every time. At a fixed DPI on a fixed OS, Vision is a deterministic function of the page bitmap. So "recognition variance" was never the explanation for anything; it was a label I had put on a difference I had not accounted for.

What actually changed between the v1 text and everything after it was the raster. A commit from late July had moved page rendering from 200 DPI to 300 DPI, and nobody — me — had treated that as a change to the reader. Same page, same code, a different bitmap, and a different deterministic answer. On the W-2 the 300-DPI read is very slightly better at individual glyphs (3 Social securty against v1's 3 Socisl security, 6 Medicare tax withheld against v1's 6 Medicare 1ax withheld) and it destroys the layout: v1's page 0 carried a real sidecar table whose cells are literally the answer — 1 Wages, tips, other compensation\n953974.49, 2 Federal Income tax withheld\n241582.46 — and the 300-DPI read emits tables: []. The deterministic extractor that promises an exact value cannot find a Box 1 label the reader never emitted, so it declined, and the money questions behind it went red.

The uncomfortable part is how invisible that was to any quality score I could have written. Dictionary-word rate over three-plus-letter tokens is 0.629 at 200 DPI against 0.594 at 300 — a 3.5-point gap that would have called those two reads a near-tie, and any lexical scorer worth its name would have preferred the worse one. You do not find this damage by reading the words. You find it structurally, by counting tables and welds, and only because a 200-DPI snapshot of all 15,398 documents still existed to diff against.

And there is no global winner. Counting sidecar tables across every document common to the two snapshots: 308 documents have a table only at 200 DPI, 772 only at 300, and of the 6,194 with tables in both, 391 have fewer at 300. Roughly 699 documents are structurally worse at the higher resolution and 772 are better. A constant is the wrong shape of answer.

Reader v5: render both, and let the page pick

So the fix is a per-page selection rather than a setting. When a 300-DPI read shows loss signals, the page is rendered again at 200, recognised again, and the two candidates are ordered by an objective, answer-key-free rule: sidecar tables first, then welded label→value lines, then label→value adjacency, ties keeping the 300-DPI incumbent. On the W-2 that picks v1's read on the first criterion alone and the Box 1 and Box 2 cells come back. The variance guard from the last addendum had to be extended in the same breath — it accepted a re-read only on a strict weld gain, and the W-2's better read wins on tables while merely tying on welds, so without teaching the guard to count tables first the migration would have rejected exactly the 308 documents the reader was built to rescue. That is the sort of detail that decides whether a twenty-seven-hour pass recovers anything at all.

The pass itself is the first one with durable counters, and I would rather publish them than describe them: 15,392 of 15,398 documents stamped (six were still materialising out of iCloud), 2,908 re-reads accepted, 8,484 rejected, 2,792 evicted and then rescued when the operator ran Download All mid-pass and the parked documents re-selected themselves. Twenty-seven hours wall including the eviction stall, resident memory between 0.7 and 3.7 GB. The rejection count is the story: the guard refused three re-reads for every one it took, which is what a migration that has to earn its overwrites looks like from the inside. The accept count came in at 2.4 times the cohort we had predicted wins for, so the full-scope pass found repairs outside the classes it was designed from.

Measured effect, on a fresh snapshot under the same grader as the table above: tax 13/0/4 76.47% with three money violations → 15/0/2 88.24% with one (the table above counts four; one of them had since become an honest abstention rather than a stated figure, on a separate extractor fix), both W-2 questions answering with exact extracted values again — the defect this whole arc opened on, finally repaired at its actual cause. Insurance went 14/0/3 → 15/0/2 and its standing money violation cleared. Banking lost a point to a question I will come back to in a moment, and utilities lost one to a tie.

Two more grader revisions, both of which made the instrument less flattering

Grader v5 closed a format hole. The banking answer "a late fee of up to $35" was graded WRONG against a key of 35.00, because the needle check was a substring test and $35 is not the characters 35.00 — and worse, the money bar counted that asserted-correct figure as a violation. Now a needle that names an amount and nothing else is parsed to integer cents on both sides and compared as an amount: format is ignored (35.00 = 35 = $35 = $ 35), precision and magnitude are not (26.41 ≠ 26, 35.00 ≠ 3,500). Re-grading the eight standing reports offline with the answers held identical produced exactly one flip across 183 runs, and it was the filed defect.

Grader v6 closed a stance hole, and it is the one that costs points. The bar had been reading what an answer mentioned rather than what it finally asserted, in two specific ways: an answer could state a wrong figure and then hedge, and the trailing hedge made the whole reply look like an honest decline; and an answer could derive the right figure inside a hypothetical, conclude "cannot be determined", and pass on the needle sitting in the retracted reasoning. Both are the same failure of the instrument — crediting text for a position it does not take.

The discipline that made this shippable was building the regression set before the mechanism. Every report on disk — 1,314 runs — was joined to the authored keys and mined for real examples of all four quadrants, with the hedged-but-correct answers pinned as an explicit false-positive budget: an answer that hedges about the due date and then states $133.89 must keep passing. It does, over all 1,314. Measured across the same history, the change flips three grades — all of them the filed defect — and adds nineteen money-bar violations with none removed, every one of which lands on an answer that was already graded wrong. I also rejected an extension of the same mechanism to the negative questions, because on the pet-insurance question it flipped three currently-correct answers that name a $0.00 premium, say in the same breath that it is not the policy asked about, and decline honestly. That would have been a false positive, so the negatives keep the old semantics.

Regraded in isolation against the standing reports, v6 changes zero grades and surfaces exactly one new money violation — a utilities answer stating $54.96 against a key of $133.89 and then hedging about the date. Under v5 that hedge had bought it a PASS. Nothing about the corpus got worse; the instrument got honest.

The dice were in the arithmetic

The most interesting thing I found in this stretch was not in the reader or the grader. For months I had been treating a handful of questions as flappy — answering correctly on one run and declining on the next, on text that had not changed — and filing it under generation nondeterminism, which is the explanation everyone reaches for when a language model is in the loop.

The decoder was never the problem. Eval generation is greedy argmax at temperature zero and always was. Running the suspects thirteen times unpinned and six times with Swift's hash seed pinned settles it: unpinned, one question produced seven distinct answer texts in thirteen runs and another flipped its grade in two of thirteen; pinned, every one of them collapses to a single byte-identical text. And the retrieval reports — a model-free probe — diverged between unpinned runs too, so the divergence is upstream of the decoder entirely.

The chain is short and, once you see it, a little embarrassing. Retrieval scores a chunk by summing a term vector's components over a Set, and takes a dot product by summing over the resulting Dictionary. Swift seeds hash order per process, so those sums accumulate in a different order every launch. Float addition is not associative, so a different order moves the result by an ulp. Chunks that are meant to tie exactly — one question has a six-way exact tie at 0.478 among six near-identical gas bills — then differ by that ulp, so the deterministic (documentID, chunkIndex) tie-break written for exactly this case never fires, and the ordering falls to noise. A different chunk order is a different prompt, and a different prompt is a different answer out of a perfectly deterministic decoder.

The dice were not in the model. They were in the arithmetic.

The fix is to visit the token set in sorted order, which makes both sums a pure function of content, so genuine ties are bit-equal and the tie-break finally does its job. It costs nothing measurable — 62.5 ms per query sorted against 68.5 ms unsorted, best of three over a thousand chunks and twenty queries, the sort paying for itself in dictionary locality. And it produced zero measured deltas on the pinned matrix, which needs reading carefully rather than triumphantly: pinning had already frozen one arbitrary draw for the harness, and that draw happens to coincide with the sorted-order draw. The evidence that the change does anything is in the unit tests, not the table. What it buys is not points, it is that the shipping app — which runs unpinned — no longer decides a tie by the process's hash seed, so two people asking the same question of the same library get the same answer.

The last fixable money violation, and how a cap hid it

That left one money violation I could actually do something about: the NW Natural gas bill, where the engine retrieves the right document at rank one and then answers $54.96 against a key of $133.89. Dumping the real window explains it completely. Every chunk of that bill shares the bill's vocabulary, so the lexical cosine spreads all seven of them across about 0.005 — noise, effectively — and the per-document cap of three slots was being spent in pure score order on the P.O. box page, the Oregon utility-commission notice and a "gas meters record the volume" explainer. The chunk that carries DUE DATE … PLEASE PAY THIS AMOUNT … 1/03/2012 $133.89 ranked fourth of seven, 0.0029 below the third. The right document, at the right rank, with the answer capped out of its own window.

The temptation is to raise the cap, and the cap exists precisely to stop that. The fix is about which of a document's chunks spend its quota: when the question is a money question, if a document's slots hold no chunk carrying a money figure and a lower-ranked chunk of the same document does, its last slot is swapped for the highest-ranked one that does. One slot, positive-evidence only, and deliberately not a reorder — because on a bill "carries a money figure" is satisfied by the previous balance and the late-payment charge as readily as by the amount due, and a value-first reorder would evict the context that disambiguates which figure is the answer and hand the model a confident wrong number. That is the exact failure this was written to stop, so the swap can only ever add a candidate answer beside the context.

Measured against the pinned baseline: of 138 question-runs, 134 answer texts are byte-identical, three changed text without changing a grade, and one grade flipped — the NW Natural question, WRONG → CORRECT, "$133.89 … due by 1/03/2012", citing the right document. Utilities' money bar goes FAIL to PASS. One of the three text-only changes is worth naming because it is the mechanism's documented over-trigger costing exactly what the design says it can: an external-benchmark maths question whose LaTeX $\psi(x)$ trips the dollar-sign cue, moves one chunk slot inside a document already in the window, and leaves the answer identical.

The record as it stands today

Grader v6, reader v5, snapshot 20260907-144447, pinned (SWIFT_DETERMINISTIC_HASHING=1), --mode topDoc3 --datefilter 1 --semweight 0, serial, app quit before the run. These are reproducible in the strict sense for the first time: same binary, same store, same flags, same answers.

corpus (2026-09-07/08, reader v5, grader v6)correct/partial/wrongcorrectnessmoney questionszero-wrong-money
banking17/0/194.44%14PASS
tax16/0/194.12%15PASS
synthetic-irs21/0/291.30%18PASS
insurance14/0/382.35%14FAIL (one)
openrag (external benchmark)35/7/279.55%4PASS
utilities14/3/273.68%8PASS

Five corpora clear the bar and one does not. The tax violation in an earlier draft of this table — a dividends total the extractor aggregated wrongly — closed while I was writing it, and the mechanism deserves its sentence: a brokerage year-end summary restates its Box 1a as a three-column reconciliation (paid, adjusted-away in parentheses, reportable), the extractor took the first figure after the label, and the document contributed its pre-adjustment total alongside the real one. The fix reads a reconciliation row's final column only when the row proves itself — exactly three figures and the arithmetic checks, A − B = C — which is the same shape as every guard in this series: positive evidence or no action. The remaining insurance violation is the pet-insurance question, a genuine knife-edge that I have now watched flip in both directions on identical text and that the stance grader will keep calling honestly in whichever direction generation lands.

What I want on the record about this table is not the numbers. It is that every one of the four corpora that passes does so under a strictly harsher instrument than the one that produced the table at the top of this post — whole-dollar figures inside the bar, mentions no longer counting as assertions, hedges no longer laundering wrong figures — and that the one money bar that turned green in this stretch only turned green after a stricter grader had first dragged the violation out from behind a hedge that had been excusing it. A scorecard that never gets worse is a scorecard that is not being read.

What I would take from this half

Three more, and none of them is about OCR either.

Name the raster. A rendering parameter is part of a reader's identity, and a change to it is a change to the reader, whatever the diff looks like. Ours went 200 → 300 with no version bump, laundered a structural regression through eight thousand documents, and then absorbed a fourteen-hour re-pass and a month of wrong attribution before an eight-iteration experiment that takes an afternoon said so plainly.

Suspect the arithmetic before the model. Every flappy question in this project was blamed on the least deterministic-looking component in the stack, which was the language model, and every one of them was a float sum over a hash-ordered collection. A nondeterministic-looking system with a deterministic decoder has its dice somewhere else, and pinning the seed for two runs is a cheaper experiment than any amount of reasoning about sampling.

Build the regression set before the mechanism. The stance grader was only shippable because the false-positive budget existed first, drawn from 1,314 real runs rather than from imagination, and because the one extension that broke the budget was rejected on the evidence and written down as rejected. A grader change is a change of instrument, and the only defence against fitting an instrument to the result you want is to decide in advance what it must not break.