Tools
The Verification Tax Meter
Generating five hundred landing pages in forty-five minutes does not eliminate the labour of writing. It moves it, in bulk, onto whoever has to establish that the pages are true — and that person’s reading speed has not improved since Gutenberg. This meter puts the two costs on the same ledger: set a volume, a generation speed and an assurance tier, and it computes what the batch actually costs in human hours, where net velocity peaks, and where it starts going backwards.
The Verification Tax Meter
v1.0.0Raw gen time
24 min
GPU time. The only number that reaches the slide deck.
Forensic audit
10 hrs
First-pass de-camouflaging at 1,200 words/hr, before fatigue.
Redline pass
3.3 hrs
Rewriting the 25% of the draft that needs it.
Editorial debt
+7.4 hrs
The N² term: hours lost purely because the pass ran past one sitting.
Recall exposure
9.3 hrs
6.2 of 14.4 fabrications get through, at 1.5 hr each.
Total human hours
31 hrs
Against 27 hrs to write the same words from scratch.
Net shipped velocity
390 w/hr
The generation speed alone implied 30,000 w/hr.
Versus writing it
1.15× slower
The machine is behind a human writing it by hand.
Directorial verdict
A drag on the studio.
This pipeline is slower than writing the same volume from scratch. The generation step still finishes in minutes, which is why the loss is invisible on the dashboard the decision was made from.
Payday loan rate: 7,589% — you borrowed 24 min of apparent saving and repaid 31 hrs of human attention. Net: 390 words/hr.
Nobody actually performs this audit. A pass this size demands 17 hrs of unbroken forensic attention — past 8 blocks the fatigue curve is describing what the audit would cost, not predicting that anyone sits through it. The real-world outcome above this line is not a slower audit. It is an abandoned one, which is how five hundred unvetted pages reach production on a Friday afternoon.
Show this chart as a table
| Pass size | Net velocity | vs. by hand |
|---|---|---|
| 250 words | 350 w/hr | 0.78× |
| 500 words | 441 w/hr | 0.98× |
| 1,000 words | 504 w/hr | 1.12× |
| 2,500 words | 521 w/hr | 1.16× |
| 5,000 words | 484 w/hr | 1.08× |
| 10,000 words | 413 w/hr | 0.92× |
| 25,000 words | 286 w/hr | 0.64× |
| 50,000 words | 195 w/hr | 0.43× |
13.3 sittings of unbroken attention
Attention degrades after 45 minutes. Detection has fallen to 57% averaged across this pass, from 72% while fresh.
390 w/hr net
Measured as shipped, publication-grade words per human hour across generation, audit, redline and cleanup — not GPU time. Writing it by hand runs at 450 w/hr.
Worth repairing
Cleanup on what escapes still costs less than starting over. The draft is a foundation, not a crack.
The same volume, batched at the crossover
Net velocity peaks at 523 w/hr for a pass of 1,927 words. Split the same 12,000 words into 7 passes of that size, with a real break between them, and the job costs 23 hrs instead of 31 hrs — 7.8 hrs back, from nothing but where you put the pauses.
The block ledger — where the N² term actually comes from
The audit is not one long smooth act. It is a sequence of 45-minute sittings, and each one is slower and less accurate than the one before it. Linear cost would be a flat column of 45s. The rise down this column is the editorial debt; the fall in the right-hand column is why the fabrications that slip through are the ones near the end.
| Sitting | Minutes | Fabrications caught |
|---|---|---|
| 1 | 45 | 72% |
| 2 | 50 | 69% |
| 3 | 56 | 67% |
| 4 | 61 | 64% |
| 5 | 67 | 62% |
| 6 | 72 | 60% |
| 7 | 77 | 58% |
| 8 | 83 | 56% |
| 9 | 88 | 54% |
| 10 | 94 | 52% |
| 11 | 99 | 50% |
| 12 | 104 | 48% |
… and 2 more sittings after these.
How this is computed — every constant, and where it came from
Two coefficients set the shape of the curve, and both are calibrated against sentences in Chapter 03 rather than tuned to make the output dramatic. Fatigue (λ = 0.12) comes from “fifty documents take two hundred times the auditing effort” of one: sitting k costs 1 + λ(k−1) times sitting one, which over fifty sittings averages out at 3.94×. Detection decay (μ = 0.037) comes from “by document twenty, subtle fabrications slip through unnoticed”: detection at sitting 20 is half what it was at sitting one. The 45-minute sitting and the 667 words/hour mission-critical audit rate are the chapter’s too — 500 words per 45 minutes.
Per-pass overhead is 18 minutes: tightening the brief, reloading context, and taking the break that actually resets the clock. It is what stops the answer being “generate one word at a time,” and it is why the crossover is a peak rather than a slope.
One number here is not computed, and cannot be. Fabrication density is not a property of the audit, the tier, or anything this page can inspect. It is a property of your model, your brief, and how far your subject sits from well-trodden training data — and no amount of reading the output tells you which case you are in. Everything downstream of it is arithmetic; this is the assumption. 1.2 per 1,000 words is a working estimate for the chapter’s scenario. Move it to whatever your own audit logs say.
| Tier | Audit rate | Redline | Caught (fresh) | Cost per escape | By hand |
|---|---|---|---|---|---|
| Low | 2,500 w/hr | 10% @ 1,800 w/hr | 55% | 0.15 hr | 900 w/hr |
| Medium | 1,200 w/hr | 25% @ 900 w/hr | 72% | 1.5 hr | 450 w/hr |
| High | 667 w/hr | 40% @ 500 w/hr | 88% | 45 hr | 250 w/hr |
Audit rate is standard copyediting throughput (~1,200 words/hour). Hand-writing rate is ~450 finished words/hour, the usual figure for a professional writer producing publication-grade copy including their own inline checking.
The shape of the curve is the finding, not the individual numbers. Generation scales linearly and cheaply: a thousand product briefs cost a thousand times almost nothing. Verification does not. It carries a quadratic term, because a human auditing synthetic prose gets slower and less accurate the longer they do it — and the smooth syntactic surface of the text is itself the sedative. Add a linear benefit to a quadratic cost and you get a peak: a batch size past which generating more raw material reduces the amount of finished work that leaves the building.
The meter finds that peak independently, and it lands where the book says it should. At the public-editorial tier the crossover comes out at roughly two thousand words — close to the 1,500-word batch ceiling the chapter recommends, arrived at from a completely different direction. Neither number was fed to the other. The two coefficients that set the curve are calibrated against sentences in the chapter itself (“fifty documents take two hundred times the auditing effort,” “by document twenty, subtle fabrications slip through”), and every constant is listed under “How this is computed” so you can disagree with a number rather than with a feeling.
One input is an assumption and is labelled as one. How many fabrications per thousand words your model actually produces is not a property of the audit or of the assurance tier. It depends on the model, the brief, and how far your subject sits from well-trodden training data — and nothing in the text reveals which case you are in, which is the whole problem. Everything downstream of that number is arithmetic. The number itself is a slider, defaulted to a working estimate, and you should move it to whatever your own audit logs say.
Push the volume fader far enough and the meter stops predicting. Past about six hours of unbroken forensic attention, the fatigue curve is describing what an audit would cost rather than claiming anyone performs it. The honest reading above that line is not a slower audit. It is an abandoned one — which is exactly how a batch of unvetted pages ends up indexed by Google before anyone has read the forty-seventh.
This is the third of three tools built on the same argument. The Median Dial shows why unconstrained output reads like nobody in particular. The Hallucination Lie Detector shows why you cannot tell when it is wrong. This one prices what those two facts cost you per batch.