Tools
The Hallucination Lie Detector
A model has no internal sense of doubt. It cannot stutter, hedge, or slide a memo across the desk with an apologetic shrug, because it has no state that separates what it knows from what it is completing. So it states an invented case citation in exactly the same unhurried voice it uses for a real one. Below are four passages. Read the first three, flag the sentences you would not sign your name to, then turn on the lens.
Brief: write a production-ready TypeScript webhook handler for Stripe charge failures, with automatic retries and exponential backoff.
Computed live from the text above, in your browser.
Fact Verification Score
Hidden until you have made your own call. Nothing on this page can compute it from the text.
The two gauges are deliberately not the same kind of number, and the difference is the entire point. The Surface Fluency Index really is computed in your browser from the words on screen: density of citation markers (years, section numbers, ratios, proper-noun runs), Latinate nominalisation, subordinate-clause texture, and — weighted highest — the total absence of hedging language. Nothing in that list has any access to whether a sentence is true. That is not a limitation of this page. It is a limitation of reading.
The Fact Verification Score is not computed at all, because it cannot be. Every claim in every passage was checked by hand against a named outside source, and those sources are listed in the bill of materials under each passage so you can go and disagree with them. Any tool that claimed to detect fabrication from the text alone would be committing the exact error it was built to expose.
The fourth passage is the control, and it is the reason to trust the other three. It says only true things, and it scores 56 on fluency against 100 on verification — the gap running the other way. Every hedge-marked phrase in it (“I checked the criterion text rather than trusting my memory,” “I could not find any deadline”) costs it fluency while costing it nothing in accuracy. Polish and proof are two axes, not one dial. A generative model is pinned to the top of the first axis regardless of where it sits on the second.
Three of these passages are modelled on failures that actually recur: a package name that does not exist on npm, a real meta-analysis quietly transplanted into a domain it never studied, and a regulatory sub-clause invented to give a deadline some authority. In each case the fabricated part is more specific than the true part around it, which is the useful tell — hyper-specificity is where a probabilistic model does its dreaming.
If you missed most of them, that is the normal result and not a comment on your attention. The whole apparatus of judging writing by how it reads was built in a world where polish was expensive. Pair this with the Median Dial for the other half of the problem: that tool shows why unconstrained output reads like nobody in particular, this one shows why you cannot tell when it is wrong. Then take the Verification Tax Meter to what those two facts cost you per batch, in hours.