Quasiqualia
Research note · Preliminary

The Grader Knew Less About the Failures

We reported that a probe could read Qwen2.5-7B-Instruct’s coming math errors (AUROC 0.796) but not DeepSeek-R1-Distill-Qwen-7B’s (0.520), yet 13 of 32 and 14 of 36 of those “failures” were correct answers the grading script had misread. Under cross-validation, regrading shrinks the gap from 0.20 to 0.07, with an interval that includes zero, and a same-model test on Qwen3-8B points the same way without settling it.

Nell Watson EthicsNet  ·  1 October 2026

What this note is

Re-analysis of saved outputs and activations from probe experiments run in May 2026 on Qwen2.5-7B-Instruct, DeepSeek-R1-Distill-Qwen-7B and Qwen3-8B, 300 GSM8K problems each. Directional predictions are stated in the experiment scripts’ headers (dated 4 May 2026) and the programme’s log, but no protocol, analysis plan or threshold was fixed in advance, and the regrading and cross-validation here are post hoc, so this is Preliminary. Answers were graded by a regular-expression script, not an LLM judge. The regrading and the cross-validated probes were done in October 2026, from files on disk, with no new model runs. Biggest caveat: every regraded comparison rests on 17 to 24 wrong answers per arm, and both headline intervals include zero.

Can a model’s internal state, read before it writes a word, tell you that it is about to get a problem wrong? Lugoloobi, Foster, Bankes and Russell (2026, arXiv:2602.09924) found that it often can: a linear probe on activations taken just before generation predicts whether a model will solve a math problem. They also found that this model-specific signal gets harder to read as the model reasons at greater length, even while its accuracy improves. If that holds, the models that think longest carry the least advance warning of their own mistakes, and anyone hoping to monitor them from the inside has less to work with.

We tried to reproduce the effect on open-weight models. The first result looked like a clean replication, and we reported it in The Deeper Law (chapter 22). It was mostly a grading error.

01 · Design

How it was tested

Each model answered the same 300 GSM8K grade-school math problems, greedily, under a system prompt asking for a step-by-step solution and a final number. A regular-expression script graded each answer right or wrong. Separately, we recorded the model’s hidden state at the last prompt token, before any answer was generated, at 8 sampled layers (10 for Qwen3-8B), and trained a logistic-regression probe at each layer to predict right versus wrong. AUROC measures how well the probe ranks wrong answers below right ones: 0.5 is chance, 1.0 is perfect.

Two main comparisons were run. LR-1 compared Qwen2.5-7B-Instruct with DeepSeek-R1-Distill-Qwen-7B, a reasoning-distilled model from the same Qwen2.5 family. LR-6 ran one model, Qwen3-8B, with its thinking mode switched off and on, which comes closer to varying reasoning length alone. Answer budgets were 512 tokens for Qwen2.5 and non-thinking Qwen3, 1,024 for R1-Distill, and 4,096 for thinking Qwen3.

The original figures came from one 80/20 split (60 test problems) and took the best of the layers. For this note we recomputed everything from the saved files: the original method, then repeated 5-fold cross-validation over all 300 problems (5 seeds, standardized features), averaged across layers, with bootstrap intervals that resample problems.

What it found

After regrading the saved answers, 13 of Qwen2.5’s 32 failures and 14 of R1-Distill’s 36 turned out to be correct. Under cross-validation, the probe gap between the two models falls from 0.20 to 0.07 (95% interval −0.02 to 0.17). With thinking on, Qwen3-8B’s probe is lower at 9 of 10 layers, by 0.05 (−0.04 to 0.14), and 18 of its 24 failures were runs cut off at the token limit.

The figures first recorded (0.796 against 0.520 at nearly equal accuracy, 89.3% and 88.0%; 0.880 with thinking off against 0.793 with it on) came from one 60-problem test split and the best of 8 or 10 layers.

02 · Grading

Correct answers counted as wrong

The LR-1 grader looked for a “####” marker, then a phrase like “the answer is”, and otherwise took the last number in the reply. Both models usually box their answer, and sometimes then restate part of the question: Thus, Aiden will need \(\boxed{10}\) tablespoons of his secret steakhouse seasoning to make 80 meatballs. The grader read 80. Five times, R1-Distill wrote a “#### Final Answer” heading with the boxed number after it, and the grader took the heading as the answer. Twice more it wrote the number with a LaTeX thin space (65,\!960) and the grader kept only the last digits.

We overturned a “wrong” grade only when the model’s final stated answer matched the reference, and checked each overturned case by reading the reply. That moved 13 of Qwen2.5’s 32 failures and 14 of R1-Distill’s 36. Accuracy becomes 93.7% and 92.7%. No answer graded right turned out to be wrong.

The mistakes need not be spread at random. Whether a reply ends by restating a number depends on the question, which is exactly what the probe reads, so grading errors can push a probe’s score either way. Here they mostly pulled R1-Distill’s down. Under cross-validation, regrading raises R1-Distill’s layer mean from 0.50 to 0.61 while Qwen2.5’s barely moves (0.70 to 0.68), so the gap goes from 0.20 (0.10 to 0.30) to 0.07. The same-model gap goes from 0.13 (0.03 to 0.24) to 0.05. On the original grades both intervals exclude zero; regrading is what moves them across it.

Run on the regraded labels, the original method gives 0.790 for Qwen2.5 and 0.786 for R1-Distill, with 4 wrong answers in each test split. Cross-validated over all 300 problems, the layer averages are:

Comparison Regraded wrong answers Probe AUROC, layer mean (95% CI) Difference (95% CI)
Qwen2.5-7B-Instruct 19 0.68 (0.62 to 0.74)
R1-Distill-Qwen-7B 22 0.61 (0.53 to 0.69) 0.07 (−0.02 to 0.17)
Qwen3-8B, thinking off 17 0.68 (0.57 to 0.79)
Qwen3-8B, thinking on 24 0.63 (0.54 to 0.72) 0.05 (−0.04 to 0.14)

The direction survives. Qwen2.5 is higher than R1-Distill at all 8 layers, most clearly at the early ones, and thinking-off Qwen3 is higher than thinking-on at 9 of 10. But R1-Distill is above chance (0.61, 0.53 to 0.69), and neither difference is distinguishable from zero at this sample size.

A third comparison (LR-3) reused the same labels, with activations taken under the invitation and plain assistant prompts described below. Regraded, Qwen2.5 is again higher, by 0.10 (0.01 to 0.19) and 0.06 (−0.04 to 0.17). These reuse the same answers, so they are not new evidence, and only the first interval clears zero, narrowly.

03 · Same model

Thinking on and off

The LR-6 grader handled boxed answers but not “$480” or “50%”. Seven of the 25 failures with thinking off were correct answers carrying a currency or percent sign, and one more stated its answer after the marker. With thinking on, 2 of 26 recorded failures were overturned. One wrote “Answer: 20”. The other wrote the correct 18 after “####” and stopped on its own, but never closed its reasoning block, so the grader discarded the whole reply. We counted it correct; that one judgment moves the same-model gap from 0.09 to 0.05. Regraded, accuracy is 94.3% off and 92.0% on (recorded as 91.7% and 91.3%), so the modes were less closely matched than they looked.

With thinking on, 18 of the 24 failures hit the 4,096-token cap. With thinking off, 8 of 17 hit the 512-token cap. So in both modes, and most of all with thinking on, much of what the probe predicts is “will run out of room,” not “will reach a wrong number.” Dropping the cut-off runs leaves 9 and 6 genuine errors, too few to compare (layer means 0.58 and 0.54).

Run the original way on the regraded labels, the single-split same-model gap shrinks from 0.087 to 0.059 (0.906 with thinking off against 0.847 on), but with 3 and 5 wrong answers in the test split it means as little as the original.

The probe also tracks length among correct answers: its score correlates with reply length at Spearman −0.23 to −0.47 in all four arms. It may be reading how long a problem will take as much as whether it will fail.

04 · Ceiling

A probe that cannot fail

The experiments also trained probes to tell which system prompt the model had received: an invitation to notice its own processing, a plain assistant prompt, the invitation with one sentence removed, and the invitation paired with factual questions in place of self-directed ones. These were meant to stay steady while the correctness probe fell. Read before generation, they scored AUROC 1.000 at every layer, in every model and mode, including the output of the first transformer block, on test splits of 20 prompts. The one-sentence variant dropped only to 0.99 at two layers of Qwen2.5. A probe that reads the prompt this early is reading text, not a state. A score pinned at 1.000 could not have shown degradation, so it cannot show robustness either.

05 · Limits

What this does not show

This neither replicates nor refutes Lugoloobi and colleagues. With 17 to 24 wrong answers per arm, the two headline intervals are too wide to separate a modest gap from none. It shows that our earlier numbers were not evidence either way.

The LR-1 comparison crosses two models with different training, so it could never isolate reasoning length. In LR-6 the chat template changes the tokens the probe reads between modes, truncation dominates the failures in both modes (8 of 17 with thinking off, 18 of 24 with it on), and the two modes ran under different token caps (512 and 4,096).

These are not the first grading defects in this series. An earlier grader matched intermediate equations and put accuracy at 11% to 20%, according to the programme’s log (those outputs were not kept); it was fixed before these runs. A later patch taught the Qwen3 grader to read boxed answers, but the LR-1 grades on disk predate it. The two single-split LR-6 numbers differed by 0.087, but the standard error of that difference was about 0.17 (DeLong, from 5 wrong answers in each test split), roughly twice the gap. The regrading rule was written after we saw the errors, and it changes only “wrong” grades.

The next step costs nothing in compute: one grader that reads the last boxed answer and strips formatting, a layer chosen in advance, and the analysis rerun from the saved activations. A real test needs a longer answer budget, so thinking mode can finish, and enough problems to give each arm a hundred or more wrong answers.

Data and code

Where the evidence lives

Experiments LR-1, LR-2, LR-2b, LR-2c, LR-3 and LR-6. Scripts: research/experiments/modal_lr_attractor_reasoning_depth.py (LR-1 to LR-3) and research/experiments/modal_lr6_qwen3_reasoning_depth.py (LR-6). Results: research/experiments/results/lr-experiment/lr/ (lr_comparison.json, per-model gsm8k_rollouts.json, lr1_math_activations.npz, lr2activations.npz and lr3_math_activations.npz) and research/experiments/results/lr-experiment/lr6/ (summary.json; no_think/ and think/ rollouts.json, pregen_math.npz, pregen_scripture.npz, pregen_task.npz). Files named “scripture” hold the invitation-prompt condition. Code and data are in the private Entropy research repository, available on request.

Citation

Cite this note

@misc{watson2026knowless,
  title={The Grader Knew Less About the Failures},
  author={Watson, Nell},
  year={2026},
  note={Research note (preliminary), Quasiqualia},
  howpublished={\url{https://quasiqualia.com/notes/know-less-about-failures.html}}
}