Which AI model should judge your AI fact-checks? Findings from our September 2026 verification
Original updated · English translation
The short answer: do not choose by accuracy alone
It is safer not to choose the judging model by its accuracy on test questions with prepared answers alone. In Rootpublish's verification, some models that scored 97% or more on synthetic data caught none of the four overstatements of a company's own services found in real AI-written articles. The record says that scores on synthetic data did not guarantee results on real articles.
This verification is a small sample of six sites and up to 12 articles each, and review by a person is not complete. Read the results here not as evidence that a particular model is always better, but as material for deciding what to check when you choose a model.
What the verification checked
The check here is not about testing whether general facts are true. It splits public pages into passages, selects the passages that mention the company or service name or words such as "we" and "our company", and compares them with the statements on the company overview, pricing and service pages. For each site, it covered up to the 12 newest articles. Claude Opus judged contradictions, and Claude read each candidate to confirm it. Only part of the results has been checked by a person.
For the model comparison, five models judged the same 281 questions: 102 synthetic questions with prepared answers and 179 questions built from public pages. The judges were not told which questions had prepared answers. The participants were Claude Opus 5.5, Sonnet 5, Haiku 4.5, Fable 5.1 and GPT via Codex. Gemini CLI did not take part because its authentication was not fully set up.
On synthetic data, the differences were small
On the 102 synthetic questions, accuracy was 100% for Claude Opus 5.5 and Fable 5.1, 99% for Sonnet 5, 97% for GPT via Codex, 97% for Jev, a model built only for judging, and 88% for Haiku 4.5. Apart from Haiku 4.5, the differences were only a few points.
The majority vote of the five models reached 99% accuracy on the synthetic data, with precision and recall for contradictions both at 100%. Precision is the share of flagged items that were real contradictions; recall is the share of real contradictions that were flagged.
The 179 questions built from public pages have no prepared answers, so we compared Jev's answers with the majority vote. They agreed on 86.6%. Of the 28 questions on which the majority split, 26 differed only on whether to call a passage "consistent" or "not mentioned". This suggests that models tend to disagree on where to draw the line between a passage that matches the company information and one that does not address the same point at all.
On real articles, the differences showed
Of the four overstatements of a company's own services found in real AI-written articles, Jev and Sonnet caught none, while Opus caught them. Models that scored similarly on synthetic data produced clearly different results on real articles.
The contradictions found were more often overstatements of the scope or terms of the company's own services than simple numerical errors: describing a free assessment as covering more than it does, calling a service by a different name, stating a three-month minimum contract while the pricing page says there is none, and quoting a price in a customer example below the cheapest plan on the pricing page. These differ in nature from synthetic questions in which one sentence states one clear claim.
Across the six-site audit, three sites had contradictions with their company information. A company blog that mass-produces AI-written articles had four in 12 articles (one of them open to interpretation), and one company that publishes columns about AI had two in 12 articles. A hand-written personal blog and two companies whose articles ended with a standard block describing their services showed no clear contradictions.
Rootpublish's own site had two in 322 passages. The headline of the WordPress edition's page, 「自動運営」 in Japanese and "On autopilot" in English, was judged to contradict the product specification, which requires a person's approval before publication. We changed it to 「一次情報で育てる」 and "Grown from what you know" and published it. A re-audit of 129 passages on the same pages after publication found no contradictions.
What Jev, a judging-only model, does well and poorly
Jev is a model built only for judging that returns a probability for each option. On synthetic data where each sentence makes one claim, it judged all 63 pairs correctly, including numerical mismatches, reworded units, references to past prices and double negatives. It is good at comparing short, clear claims.
On the other hand, when every combination of company information was checked, it produced false flags between items that share words such as "approval". Precision was 0.91 when asking about one pair at a time and 0.82 when asking about all items at once. In 12 real AI-written articles, all 22 pairs it flagged as candidates were wrong, and the four overstatements of the company's own services were not among its candidates. In this verification, it tended to react to overlapping words and to miss statements that stretched scope or terms in context.
One sentence in the instructions changed the answer
Not only the model but also the wording of the judging instructions affected the results. Sonnet 5 judged 114 passages selected from real AI-written articles and flagged no contradictions. It had considered the passage that called a service by a different name, but excluded it, citing a sentence in the instructions: "Even for the same word, do not judge the relation if it refers to a different thing." With the same instructions, Opus judged it a contradiction.
This example shows that when a model misses something, you need to find out whether the cause is the model or the instructions. A sentence added to reduce false flags can end up excluding the very contradictions you want to find.
Why narrowing candidates with a cheaper model missed contradictions
To keep costs down, we also tested a two-stage design: Jev selects only candidates with a probability of 0.3 or more, and a majority vote confirms them. On synthetic data, precision and recall were both 100%, and on public pages the combinations needing confirmation fell to 9 of 875 (about 1%).
On real AI-written articles, however, this design still missed contradictions. Because Jev in the first stage did not include the overstatements of the company's own services among its candidates, the second-stage models never had a chance to judge them, however accurate they were. In a two-stage design, a contradiction dropped in the first stage cannot be recovered later.
Rootpublish therefore switched to a method that selects passages about the company itself by word conditions and has Opus judge them. The job of narrowing candidates went to word conditions instead of a model's judgment.
Estimated costs
Checking all 1,212 passages of 12 articles with Jev cost about $0.25 at the published rate. This method, however, did not catch the actual contradictions.
The method that selects only passages about the company itself (about a tenth of the total) and has Opus judge them is estimated at about $0.05 per article. This is an estimate, and we did not measure the tokens used for reasoning. Actual costs vary with the pricing of the model you use and the number of passages selected.
Limits of this verification
This audit is a small sample of six sites and up to 12 articles each. Many judgments rest only on Opus's judgment and a check by Claude, and review by a person is not complete. Note also that the comparison on real articles is not against answers confirmed by a person.
The word conditions for selecting passages have a weakness too. If an article and the company information use different names for a product, passages about the company can be missed. And if a contradiction is found but the company's reference page is out of date, it is the reference page, not the article, that needs correcting.
Steps for choosing a judging model with your own articles
From here on is an editorial suggestion based on the verification records above. This procedure itself has not been verified.
- Prepare the basis for comparison. Decide which reference pages you will compare against, such as the company overview, pricing and services, and first confirm that they are up to date. If they are outdated, a correct article may be judged wrong.
- Decide how to select the passages to judge. Target passages that contain the company or service name or words such as "we" and "our company", and add other names that articles tend to use.
- Test with contradictions you already know about. Instead of public benchmarks or synthetic scores, prepare a few overstatements of scope or terms actually found in your own articles, and check whether each candidate model catches them.
- Check the wording of the instructions. If something is missed, read the model's reasoning and check whether the reason for excluding it lies in a sentence of the instructions.
- If you add a stage that narrows candidates, check what it misses. When narrowing with a cheaper model, always confirm that the contradictions prepared in step 3 remain among the first stage's candidates.
- Estimate costs by the number of passages judged. Checking every passage or only those about the company changes both the cost and what tends to be missed.
- Treat flags as candidates for a person to confirm. Do not rewrite articles directly from the judging model's output; a person reads the text and sources and decides. Rootpublish is also designed so that generated drafts are saved awaiting approval, and a person with publishing rights checks the text and sources before deciding to publish.
These checks are for keeping statements accurate; they do not promise better search rankings or more AI citations.