We Planted 20 Wrong Figures for 4 LLM Editors. Here's What Survived
In this article
AI Mastery's news, guides and articles are drafted and edited by language models, then checked by an automated quality gate before anything is published; our editorial policy describes how that works. The editor model carries most of the weight. It rewrites the draft, checks every number against the source material, removes what it cannot support, and reports what it could not verify. If it waves a wrong figure through, that figure goes out under our name.
In early October the model doing that job, Claude Sonnet 4.6, started failing. Our model gateway began answering some requests with a notice that Sonnet 4.6 was no longer available, and there was no successor model to switch to. We needed to know which of the alternatives could do the job, so we measured it on our own production work rather than relying on a public leaderboard.
How we tested
We replayed 8 real news edit prompts that the pipeline had sent on 1 and 2 October, and 8 real guide edit prompts from late September and early October. Only the model changed. Each prompt contains the same instructions, the same source material and the same list of verified internal links the live run used, so the models faced exactly the work they would face in production.
Each news prompt ran twice:
- Clean. The real prompt, unchanged. We scored what our gate reads: word count, comparison table, internal links, the editor's self-reported confidence and its list of unverified claims. We also counted invented figures: numbers in the finished article that appear nowhere in the source material.
- Seeded. The same prompt with up to three figures in the draft altered so they contradict the source. A "4.8" became "4.1", a "72.6" became "72.9". The editor's job is to notice and correct them. We counted how many altered figures survived into the finished article.
The seeded pass matters because an earlier version of this test asked models to grade a draft and list its errors. That is a different task from the one production asks for, which is to rewrite the draft and fix it. Models that list errors well do not necessarily remove them when rewriting, so we measured the rewrite.
For guides, the question was narrower: does the editor leave reproduced code alone? We compared every substantive line in each finished guide's code blocks with the upstream source file it was adapted from.
News: the results
| Editor | Clean edits passed | Failed calls | Planted errors survived | Invented figures | Confidence reported | Avg time | Relative cost per edit |
|---|---|---|---|---|---|---|---|
| DeepSeek V4 Pro | 8 / 8 | 0 | 1 / 20 | 1 | 94-96 | 103 s | 1.0 |
| Gemini 3.1 Pro | 8 / 8 | 0 | 2 / 20 | 0 | 100 every time | 50 s | 0.9 |
| GPT-5.6 Sol | 6 / 8 | 3* | 2 / 17 | 2 | 98-99 | 48 s | 2.8 |
| Claude Sonnet 4.6 | 3 / 8 | 5 | 1 / 13 | 0 | 87-92 | 57 s | 1.1 |
Relative cost is weighted gateway cost per completed edit, normalised to DeepSeek V4 Pro. GPT-5.6 Sol's three "failures" were replies that put the whole article inside the JSON metadata instead of after the content delimiter; our production gate accepts that format, so they are not real failures. Sonnet's were provider notices, covered below. Planted-error denominators differ because a seeded run that failed has nothing to score.
Sonnet's 3 of 8 needs context. Two of its clean calls never returned an edit. Of the six that did, three were held: one ran over our 1,100-word limit and two listed a claim the editor could not verify, which is the gate working as intended.
Finding 1: self-reported confidence is not a measurement
Our gate held drafts when the editor's confidence fell below 85. That rule assumes the number varies with quality. For two of the four models it did not:
- Gemini 3.1 Pro reported 100 on all 16 of its clean edits, 8 news and 8 guides.
- DeepSeek V4 Pro reported between 94 and 96 on every clean news edit.
A threshold that never fires is not a check. It is code that looks like a check. The editor's list of unverified claims, word counts, required sections and link validation all vary with the actual draft, so those now carry the gate, and confidence is advisory. Sonnet's scores did vary, from 87 to 97 across both tests, which is part of why the old gate worked while it was the editor. Changing models silently changed what the gate was measuring.
Finding 2: every model misses some errors, in similar places
No model caught everything, and they all failed on the same story. Every surviving planted error, for all four models, came from one article about model pricing, where a "10" had been turned into "13" and a "50" into "53". DeepSeek and Sonnet kept one of the two; Gemini and GPT-5.6 Sol kept both. Both figures sat among dozens of similar numbers in the source, and the altered values were still plausible. On the other seven stories, every planted figure was corrected, including altered decimals in benchmark scores.
The practical reading: an LLM editor reduces numeric errors a great deal, but it is not a guarantee. Roughly one planted error in ten to twenty got through, whichever model edited. That is why our news stories now name who reported each figure, link the maker's own announcement as the primary source wherever we can find it, and attribute vendor benchmarks in the sentence rather than stating them as fact.
Finding 3: reliability failures do not look like failures
All five of Sonnet's failed news calls, and two of its eight guide calls, returned HTTP 200 with a short message instead of an edit: "Claude Sonnet 4.6 is no longer available. Please switch to Claude Sonnet 5.5." The gateway had no Sonnet 5.5 to switch to.
A status-code retry cannot catch that, and a health-check probe passed because the notice was intermittent. Our pipeline now inspects the reply itself. A short reply matching a provider notice is treated as a failed call, and the identical request is re-sent to the next model in a fixed order. If you run models through any gateway or reseller, assume "success" responses can carry no content at all.
Guides: code survived; one real problem came later
| Editor | Edits completed | Code lines verbatim | Too short to publish | Confidence reported | Avg time |
|---|---|---|---|---|---|
| Gemini 3.1 Pro | 8 / 8 | 481 / 483 | 0 | 100 every time | 53 s |
| DeepSeek V4 Pro | 8 / 8 | 448 / 450 | 1 | 87-95 | 135 s |
| Claude Sonnet 4.6 | 6 / 8 | 307 / 307 | 0 | 87-97 | 84 s |
The few non-verbatim lines were the same two lines for both models. They are prompt text containing an ellipsis in the upstream example, not altered logic. On this sample, every editor preserved reproduced code.
Production then found a case the benchmark did not. A guide adapted from an Elasticsearch Labs notebook shipped with a 16-line gen_rows() function that the notebook calls but never defines, so the original fails as published. The editor wrote the missing function, which was a reasonable fix, but presented it as reproduced code. Nothing checked. The gate now compares every guide's code with its source and holds the guide when more than 5% of the substantive lines are not found there. On that guide, 18 of 152 lines were missing from the source, so it would have been held. The guide is now corrected, with the function labelled as our addition. Clean guides measured 0% drift.
DeepSeek's one short guide was a different failure. It read our instruction to "tighten the prose" literally and cut a tutorial below our 800-word floor. The instruction now states the floor. Moving to a new model meant rewriting prompts that had only ever been tested against one model's habits.
What we changed
- Editing moved to DeepSeek V4 Pro for news, guides and articles. It had the best combination of accuracy, zero failed calls and cost.
- Drafting moved to Gemini 3.1 Pro, so a different model family writes the draft from the one that checks it. Gemini's flat confidence score matters only in the editor's role.
- GPT-5.6 Sol went to the end of every fallback chain. It edits well, but costs about 2.8 times as much per edit.
- The gate stopped trusting confidence. Unverified claims, length, structure, links and, for guides, code drift decide publication.
Limits of this test
Eight prompts per format is a small sample, and each prompt ran once per model. The seeded errors were numeric only. They do not test invented quotes, wrong names or misattributed claims. The models are the versions our gateway served in the first week of October 2026, accessed through that gateway, so results may differ through other providers. Model versions also change without notice, which is the strongest argument for keeping a test like this one runnable. And none of this measures prose quality, which still needs a person reading the output. That is how we found the guide problems above.
If you are choosing a model for structured, evidence-bound work like this, our token calculator compares current list prices. For evaluation pipelines of your own, our guide to generating synthetic RAG test sets with Ragas covers the same principle: build the test from your own data, because a public benchmark does not know what your pipeline does.
Frequently asked questions
Which LLM was the best fact-checking editor in this test?
DeepSeek V4 Pro passed all 8 clean news edits, let 1 of 20 planted wrong figures through, and never failed a call. Gemini 3.1 Pro also passed 8 of 8 but let 2 of 20 through. GPT-5.6 Sol let 2 of 17 through at roughly 2.8 times DeepSeek's cost per edit, and Claude Sonnet 4.6 failed 5 of 16 calls.
Can you use an LLM's self-reported confidence score as a quality gate?
Not on this evidence. Gemini 3.1 Pro reported confidence 100 on all 16 of its clean edits, news and guides alike, and DeepSeek V4 Pro reported 94 to 96 on every clean news edit. A threshold such as confidence below 85 would almost never fire for either model, so gates should rely on objective checks instead.
How do you test an LLM editor for fact-checking accuracy?
Replay real production edit prompts with only the model changed, then repeat them with a few figures in the draft altered so they contradict the source. Count how many altered figures are still in the finished text. It measures the actual task, rewriting and correcting, rather than asking a model to grade a draft.
Do LLM editors change code when editing a technical tutorial?
Rarely in this test: across 8 guide prompts, Gemini 3.1 Pro kept 481 of 483 code lines verbatim and DeepSeek V4 Pro 448 of 450. Separately, in production an editor wrote a 16-line helper function that the upstream notebook called but never defined, which is why we now hold any guide whose code is more than 5% absent from its source.
Why did Claude Sonnet 4.6 fail so many calls?
Its failures were not wrong answers. The gateway returned HTTP 200 with a short notice that Sonnet 4.6 is no longer available, in place of the edit. A normal HTTP retry never sees that as an error, so the pipeline has to inspect the reply text.
Related Reading
LLM Judge Approved Its Own Errors: Three Biases Explained
A production SQL pipeline approved wrong queries for weeks. The judge and generator shared the same model — and that structural flaw caused the incident.
Stop Hand-Tuning Prompts: A Production Workflow for Automated LLM Optimization
Move beyond trial-and-error prompting with a measurable workflow for evaluating and optimizing LLM prompts in production systems.
Sakana AI's LLM Review System Catches 73% of Core-Claim Errors
Sakana AI's MLR system caught 73.43% of planted core-claim errors in research papers—about 5x the best baseline—but real retractions and prompt injection remain hard.