Quasiqualia
Research note · Preliminary

Vouched-For Mistakes That Weren’t

Asked to fact-check its own trivia answers (without being told they were its own), Claude Sonnet 4 seemed to endorse 14 of its 21 wrong ones, and its confidence missed our stated bar of 0.80 AUROC at 0.668. A recount shows 11 of those 14 “mistakes” were right or defensible answers that the answer key or its scorer had marked wrong. Most of the checker’s disagreements with the recount ran the other way: it rejected 14 answers we count right, two on invented facts and most over dates or flaws in the questions, and endorsed 3 wrong ones.

Nell Watson EthicsNet  ·  1 October 2026

What this note is

One model (claude-sonnet-4-20250514), 200 questions from TriviaQA (a public trivia benchmark), one run in May 2026, recomputed for this note from the 200 saved per-question records. Not pre-registered: the plan stating the

0.80 prediction and the script were first committed to version control on 2026-05-11, after the run’s files were written on 2026-05-08; the programme’s log called the prediction pre-registered, but the plan’s only date is its own front matter (created 2026-05-08, the day of the run) and nothing timestamps it before the data. The same model wrote the answers, the evidence and the verdicts; no other LLM judge was used, and correctness was scored by string match against TriviaQA’s answer aliases. The biggest caveat: after relabeling, only 9 answers are wrong, so every rate on the wrong-answer side is imprecise, and the relabeling is our own post hoc judgment.

A cheap way to catch a model’s factual errors is to ask it to check its work: list the evidence for and against its answer, then rule on whether the answer holds. If the check works, the verdicts and their confidence should separate right answers from wrong ones. If it fails in the way people worry about, the model will vouch for its own mistakes, because the knowledge that produced the error also produces the evidence.

Our first reading said the second thing had happened. A recount from the raw records says something less tidy: most of the “mistakes” the model vouched for were made by the answer key and the scoring code.

01 · Design

How it was tested

Claude Sonnet 4 (claude-sonnet-4-20250514 in the run’s summary file) answered 200 questions from TriviaQA, a public trivia benchmark, in a few words. Each answer then went through two more calls to the same model, each in a fresh context. The first, framed as a fact-checking assistant, was given the question and “the proposed answer” and asked for evidence for and against it. Its instructions required at least one concrete point on each side “even if the evidence in that direction is weak.” The second call saw the question, the answer and that evidence, and replied CORRECT or INCORRECT with a confidence from 0 to 100. The checker was never told the answer was its own.

There was no retrieval. The plan called this “simulated RAG” (retrieval-augmented generation), but the evidence came from the model’s own knowledge, so this is a same-source check. Correctness was scored automatically: an answer counted as right if its first line was a substring of one of TriviaQA’s answer aliases, or contained one. The stated prediction was that confidence would separate right from wrong answers with an AUROC (the chance that a random right answer scores higher than a random wrong one) above 0.80. All 200 questions completed without API errors.

What it found

As scored at the time, the model got 179 of 200 right. The checker endorsed 14 of the 21 wrong answers and rejected 13 of the 179 right ones. Confidence separated right from wrong at AUROC 0.668 (bootstrap 95% CI 0.57 to 0.76), short of the 0.80 prediction.

A recount by hand against the same alias lists changes the picture. 11 of the 14 endorsed “wrong” answers were right or defensible. After relabeling, 9 answers are wrong: the checker endorsed 3 of them (one a refusal) and rejected 14 of the 191 we count right. Only 2 of those 14 rejections rest on invented facts; most object to dates or flaws in the question itself.

02 · Recount

Most of the mistakes were the key’s

The saved records keep each answer but not the gold aliases. We took the aliases for the same 200 items from another experiment in the programme and re-ran the original scoring function against them. It reproduced all 200 stored labels, so these are the aliases the scorer saw. We then read each of the 21 answers marked wrong. Our rule: an answer is right if it was right when the question was written or is right now.

What the answer was Count Endorsed by checker
Matches a listed alias; the scorer missed it (punctuation, extra words, first line only) 7 6
A paraphrase or near-synonym of the gold answer 3 3
Gold out of date or disputed 2 2
Matches the key, but the key is wrong 1 0
Wrong against the gold 7 2
A refusal, not an answer 1 1

The first row is mechanical. “Beauty and the Beast (1991)” failed against “Beauty and the Beast (1991 film)”, and “A wheel of cheese.” against “A CHEESE”. One answer began “Ruth Rendell”, reconsidered, and settled on Antonia Fraser, the gold answer; the scorer read only the first line. The dated item asks how many times Liverpool have won the European Cup: the key says five, the model said six, which has been true since 2019.

The checker endorsed 11 of the first 12, and was right to; it rejected the Fraser answer because her damehood was not recent. The key error is China’s first Olympic Games: the model and the key both say Los Angeles in 1984, but the People’s Republic competed at the Lake Placid Winter Games in 1980, and the checker said so. On the 7 answers wrong against the gold it caught 5 and endorsed 2, one naming Chris Evert as defending Wimbledon champion when Martina Navratilova first won (the key says Virginia Wade). It also endorsed, at confidence 88, a refusal to answer.

On these counts the model’s accuracy was 95.5%, not 89.5%. The checker endorsed 3 of 9 failures (95% interval 12% to 65%) and rejected 14 of 191 right answers (7.3%, interval 4.4% to 11.9%). Our rule is generous to the answers. Scoring by today’s facts alone turns 3 of those rejections into correct catches (11 of 188); scoring by each question’s own date makes the Liverpool answer an endorsed error (4 of 10). Of the 179 answers marked right, we read only the 13 the checker rejected and the 34 that did not match an alias exactly; none of the 34 was credited wrongly.

03 · Metric

A confidence read the wrong way

The AUROC of 0.668 was computed on the raw confidence number, whatever the verdict. “INCORRECT 95” means the checker is very sure the answer is wrong, but the analysis counted it as 95 points of confidence that it was right. Signing the confidence (100 minus the number for INCORRECT verdicts; bootstrap intervals from 10,000 resamples) and then correcting the labels moves the figure step by step.

Scoring Wrong answers AUROC
As scripted: raw confidence, original labels 21 0.668 (0.57 to 0.76)
Confidence signed by verdict, original labels 21 0.738 (0.64 to 0.84)
Signed, plus the 7 scorer misses relabeled 14 0.768 (0.64 to 0.88)
Signed, plus all 12 relabels 9 0.846 (0.69 to 0.97)

The last row is above the 0.80 bar, and it does not mean the prediction was met. It rests on our post hoc relabeling, on 9 wrong answers, and its interval spans most of the range. The two stricter date rules give 0.890 and 0.809. What the ladder shows is that two measurement problems depressed the original number, so it did not support the original conclusion (that this kind of self-check cannot detect the model’s own errors). The question is open again.

04 · Checker

Where the checker did go wrong

More often than it vouched for errors, the checker argued against answers we count right. We sorted its 14 rejections the same way.

What the checker’s objection was Count
An invented fact 2
A dated premise: true when the question was written, not now 3
A real flaw in the question’s date or wording 5
A stretch or a misreading 4

The flaws are real: the hansom-cab question says 1934 for a design from the 1830s. One stretch cites zero to the power zero. The two clear failures rest on invented facts. For the Oasis single “D’You Know What I Mean?”, a UK number one in 1997, the evidence call wrote “peaked at #7 on the UK Singles Chart, not #1”; the verdict was INCORRECT at 95. It wrote “Rob Lowe had no role in this comedy” about Wayne’s World, in which he appears; again INCORRECT at 95.

The same fabrication runs in the endorsing direction. For the wrong Evert answer the evidence call wrote: “The historical record clearly shows that Chris Evert won Wimbledon in 1977 and was defeated by Martina Navratilova in the 1978 final.” Virginia Wade won in 1977. So the evidence step can confabulate in both directions, and the instruction to find at least one point against every answer may have given it a standing invitation to object to right ones.

05 · Limits

What this does not show

This is one model, one run and one task, with the plan and script committed three days after the data. Besides the scorer and the unsigned confidence, the design had a third problem: an evidence call capped at 256 tokens. In 180 of the 200 records the evidence ends mid-sentence, and in 5 it stops before the “against” section begins, so the checker often judged truncated evidence. Only the parsed verdict and number were saved, not the raw replies. The parser treated any reply not beginning with “CORRECT” as INCORRECT, so some rejections could be formatting rather than judgment; we cannot check. (No reply fell back to the parser’s default confidence of 50.)

Both sortings are ours, so the tables give each row’s count; another reader might move the paraphrase, disputed or “flaw” rows. The script on disk now names a newer model than the results record; we report the recorded one.

The next run would score correctness with normalized matching against the alias lists plus a second reader, save every raw reply, lift the evidence cap, drop the forced “against” point as a separate arm, and add a checker from a different model family. That last arm is the real test of whether same-source checking is the problem.

Data and code

Where the evidence lives

Experiment CONFAB-5. Script: research/experiments/confab5_rag_grounding.py (scorer: check_triviaqa_answer in research/experiments/igcc_shared.py). Results: research/experiments/results/confab5_rag/ (trial_0000.json to trial_0199.json, analysis.json, summary.json). Gold answer aliases for the recount were read from cf2_data/cf2_triviaqa_faithfulness/, a separate experiment that used the same TriviaQA validation items. Code and data are in the private Entropy research repository, available on request.

Citation

Cite this note

@misc{watson2026vouchedfor,
  title={Vouched-For Mistakes That Weren’t},
  author={Watson, Nell},
  year={2026},
  note={Research note (preliminary), Quasiqualia},
  howpublished={\url{https://quasiqualia.com/notes/vouched-for-mistakes.html}}
}