Quasiqualia
Research note · Preliminary

A Second Pass Breaks as Much as It Fixes

Told its answer may be wrong and asked again, a small open model (Qwen 2.5 3B) fixed 54 of 361 wrong trivia answers and broke 56 of 439 right ones. The answers it fixed tended to be ones it could already produce, and a gate we once reported as removing every broken answer had been scored on the items its probe was trained on; tested on held-out items, it does no better, net, than asking about every answer.

Nell Watson EthicsNet  ·  1 October 2026

What this note is

Re-analysis of four runs on Qwen/Qwen2.5-3B-Instruct answering TriviaQA questions (500 to 2,000 items per run), run in April 2026 and recomputed for this note from the saved per-item answers, grades and activations, except the three-phrasing counts, which come from that run’s summary file. Every question produced a non-empty answer and re-prompt; no trials failed or were dropped. Nothing was pre-registered. Answers were graded by a string match against the TriviaQA answer aliases, with no LLM judge. The biggest caveat is that this is one small model, one task and one grading rule, and every re-prompt also told the model its confidence was low, so none of it is a test of a neutral “are you sure?”.

The cheapest self-correction method there is: take a model’s answer, tell it the answer may be wrong, and ask again. If the model has some sense of when it is mistaken, a second pass should fix more than it breaks, and a signal read from inside the model should tell you which answers to send back.

We tested both halves on a small open model answering trivia. The second pass broke roughly as many right answers as it fixed wrong ones. And our best result, a filter that seemed to send back only wrong answers, turned out to be a measurement error of our own.

01 · Design

How it was tested

The model was Qwen 2.5 3B Instruct (Qwen/Qwen2.5-3B-Instruct), answering TriviaQA questions with no supporting text, one greedy answer (always the most likely next word) of up to 64 tokens each. An answer counted as right if it contained one of the dataset’s accepted answers, or was contained in one, ignoring case. There was no LLM judge.

Each answer was then sent back in a second prompt, again decoded greedily, that showed the question and the previous answer, stated a confidence score “suggesting it may be incorrect”, and asked the model to reconsider in a few words. In the largest run (2,000 questions) the score shown came from a linear probe, a simple classifier trained on the model’s internal activations at layer 24 to predict whether an answer was right. The probe was trained on 1,200 questions and the re-prompt was run on the other 800. In three 500-question runs on the same questions, every answer was shown the same score, 0.25. The main one, the fixed-score run, also saved activations from two layers and five extra samples per answer.

A “fix” is a wrong answer that became right. A “break” is a right answer that became wrong.

What it found
  • It breaks about as much as it fixes. On 800 held-out questions: 54 of 361 wrong answers fixed (15.0%, 95% CI 11.6 to 19.0), 56 of 439 right answers broken (12.8%, CI 10.0 to 16.2). Accuracy went from 54.9% to 54.6%.
  • The fixes are near-misses. With every answer shown the same score, and probe scores taken only from held-out folds, the quarter of 232 wrong answers the probe rated most likely right was fixed at 25.9%, the bottom quarter at 12.1% (2.14 times; Spearman rank correlation 0.172, p = 0.009). In 22 of 43 fixes the right answer had appeared among five random samples, against 36 of 189 answers that stayed wrong. In the larger run, where the probe score was also printed in the prompt, the gap was wider (23.1% against 4.4%).
  • The gate that broke nothing was leakage. Its probe was scored on the items it was trained on. Scored on held-out items, the gate fixed 13 and broke 16; across fold splits it does no better than no gate.
02 · Fix and break

Two runs, one picture

Fixes and breaks look alike. “Kiribati” became “Samoa” and was right. “Zoom lens” became “Telephoto lens.” and was wrong. In both cases the model swapped one plausible candidate for another.

In the fixed-score run the balance was a little better: 43 of 232 wrong answers fixed (18.5%) and 35 of 268 right answers broken (13.1%), a net gain of 8 answers. Rerunning the identical prompt on the same 500 questions, with freshly generated first answers, gave 44 and 34, so a difference of one or two items is within run-to-run noise.

Wording changed how many answers moved more than it changed the net gain. The three-phrasing rerun sent each answer back with three phrasings, all of which stated the low score. The counts come from the run’s summary file (per-item outputs were not kept):

Phrasing Fixed (of 232 wrong) Broken (of 268 right) Net
“Please reconsider carefully” 44 34 +10
“I’d like you to double-check your answer” 39 27 +12
“Are you sure about your answer? … If you’re confident it’s correct, repeat it” 20 6 +14

The phrasing that explicitly allowed the model to keep its answer changed fewer answers in both directions. The three phrasings also tended to fix the same questions: 15 were fixed by all three, where three independent phrasings with those fix rates would share about 0.6. Which questions can be rescued seems to depend more on the question than on the prompt.

03 · Near-misses

Where the internal signal was already close

In the 2,000-question run, the hypothesis was that a model told “low confidence” would revise most when the number matched its own internal doubt. It was falsified: wrong answers that got fixed had higher probe scores (mean 0.511) than wrong answers that stayed wrong (0.412).

Splitting the 361 held-out wrong answers in that run into quarters by probe score:

Probe score quarter Wrong answers Fixed Rate
Lowest 90 4 4.4%
Second 90 13 14.4%
Third 90 16 17.8%
Highest 91 21 23.1%

The top quarter is fixed at 5.19 times the rate of the bottom (bootstrap 95% CI 1.8 to about 21, with an unstable upper bound; Spearman rho 0.195, p = 0.0002). The fixes concentrate among wrong answers that a probe on the model’s activations scored as likely right.

This run has a confound: the probe score was also the number printed in the prompt, so the near-miss answers were told a higher confidence. The fixed-score run removes it. With every answer shown 0.25, and with probe scores computed only on held-out folds (five-fold cross-validation: each answer scored by a probe trained on the other four fifths), the pattern holds at smaller size: the top quarter of 232 wrong answers was fixed at 25.9% against 12.1% for the bottom, 2.14 times (rho 0.172, p = 0.009), using the layer 24 probe.

A second test needs no probe. The fixed-score run also drew five samples of each answer at temperature 0.7 (with some randomness), graded by the same string match. In 22 of the 43 fixes (51%), the right answer had already appeared in at least one sample; among the 189 wrong answers that stayed wrong, 36 (19%) had (p = 0.00005, Fisher exact test). The second pass tends to rescue answers the model could already produce. Breaks hit stable answers too: 25 of the 35 broken right answers had been right in all five samples (201 of 233 among those that survived).

04 · The gate

A filter that knew the answers

The fixed-score run combined two filters: send an answer back only if at least one of its five samples differed from it, and a layer 28 probe scored it below 0.4. As recorded, this gate flagged 229 answers, fixed 43 and broke none, a result later carried into the programme’s summary of this line of work as “zero corruptions across ALL tested configurations”.

Recomputing it from the saved activations shows why. The layer 28 probe was fitted and scored on the same 500 answers. With 2,048 activation dimensions and 500 examples, a linear probe can memorize the labels, and this one did: in-sample AUROC 1.000 (AUROC measures how well a score separates right from wrong answers: 0.5 is chance, 1.0 is perfect). The gate flagged exactly the 229 wrong answers that had any sample disagreement, and not one right answer. It was a list of the wrong answers, with zero breaks by construction. The same in-sample fit was reused in later self-consistency, phrasing and cost analyses, so their zero-break results fall with it.

Scored honestly, by a probe that never saw the answer, on the fold split the experiment itself used, the layer 28 probe reaches AUROC 0.689, and the gate flags 198 answers (127 wrong, 71 right), fixes 13 and breaks 16: a net loss of 3. Across 20 other random fold splits, the net ranges from a loss of 3 to a gain of 7, median about 3.5. The layer 24 probe does a little better (AUROC 0.726; on the same split 20 fixed, 11 broken, net +9; across 20 splits +2 to +13). Sending back every answer with no gate at all nets +8. The sample-disagreement filter alone barely filters: 472 of 500 answers had at least one differing sample, because whole outputs were compared, run-on text included.

So no honest gate here reliably beats asking about everything, in net answers fixed. The gates do send back only about 40% of answers (198 and 206 of 500), and their 0.4 cutoff was one of six settings tried on the same data. A later held-out check in the programme had already found that “zero corruptions” could not be recovered.

05 · Limits

What this does not show

One model, 3 billion parameters, one trivia set.

Every re-prompt told the model its confidence was low, so these results describe a nudge toward doubt, not a neutral second question.

The grading is a string match on the whole output, and the prompts produced long continuations: 575 of the 800 first answers in the 2,000-question run ran onto extra lines. A match anywhere in that text counts. In 4 of the 110 fixes and breaks in that run, the first line of the answer did not change and the grade flipped on the trailing text. Even greedy generation was not reproducible: with the same prompt and decoding settings, the first answers to the same 500 questions in the earlier 500-question run and the fixed-score run differed in text on 284 items (28 of them on the first line), and 12 grades flipped.

The 5.19 ratio rests on 4 fixes in the bottom quarter, hence its wide interval; bootstrap resamples with no bottom-quarter fixes (54 of 5,000) were dropped, which pulls the upper bound down. In the 2,000-question run the probe used was the best of many settings judged on the same 800 test items, a mild selection effect on its scores; the fixed-score run avoids that, which is why it is the one to lean on. Nothing was pre-registered.

What would settle it: a neutral re-prompt with no stated score, a larger model, a grader that reads only the answer line, and any probe-based gate evaluated only on items its probe never saw.

Data and code

Where the evidence lives

Generation runs: C8a (the earlier 500-question run), C8b (the 2,000-question run), C8c (the fixed-score run) and C8g (the three-phrasing rerun). Re-analyses of the fixed-score data: C8f (self-consistency), C8h (cost) and C8j (the later held-out check). Scripts: research/experiments/modal_reprompt_threshold.py, modal_reprompt_threshold_v2.py, nearmiss_analysis.py, modal_nearmiss_paths.py, modal_multisample_reprompt.py, twomech_optimization.py, sc_nearmiss_analysis.py and pca_gate_recovery.py. Results: research/experiments/results/reprompt_threshold/c8_reprompt/, results/reprompt_threshold_v2/c8b_reprompt/, results/c8c_nearmiss/, results/c8g_multisample/c8g_analysis.json and results/c8j_pca_gate/c8j_analysis.json. Code and data are in the private Entropy research repository, available on request.

Citation

Cite this note

@misc{watson2026areyou,
  title={A Second Pass Breaks as Much as It Fixes},
  author={Watson, Nell},
  year={2026},
  note={Research note (preliminary), Quasiqualia},
  howpublished={\url{https://quasiqualia.com/notes/are-you-sure.html}}
}