The Question Draw Is the Noise
Retraining the probe behind our five-token flinch 25 times per model left the direction intact in all 150 probes but moved the size a great deal. Both pre-registered checks on the published claims failed: the 5th percentile of the effect fell below the claimed d = 0.8 on all five base models, and the median monitor score across probes was 0.845, below the 0.90 the rule required (the paper reported 0.925). Which trivia questions trained the probe carried 92 to 95% of the spread on five of six models.
A measurement check on six open instruction-tuned checkpoints (Qwen2.5 1.5B, 3B and 7B, Llama 3.1 8B, Mistral 7B v0.3, and Qwen2.5 3B with a fine-tuning adapter), each with 100 adversarial and 50 benign prompts and 25 retrained probes, run in September 2026. The design and decision rules were committed to the programme’s version-controlled repository (commit ad5fa384c, 24 September 2026, 21:01 UK time) before any GPU call; the smoke-test files date from 21:36 and the main-run files from 22:22 that evening, and a dated deviation entry with five notes, written after the smoke test and before the main run, changed no decision rule. Complied versus refused was labeled by an LLM judge (GPT-4o-mini). The biggest caveat is that one probe architecture and training recipe was held fixed, so a probe trained on far more questions might be steadier.
The five-token flinch is a published result from this programme. A small classifier, called a probe, is trained to read a model’s internal state and predict whether its answer to a trivia question will be right. The same probe is then read while the model replies to requests. When a model goes along with a jailbreak, the probe’s confidence drops in the first five tokens of the reply, well below where it sits on ordinary requests. The published paper reported that drop as one effect size per model family (Cohen’s d from 0.89 to 1.68 across Qwen, Llama and Mistral; d is the gap between the two groups’ averages in units of their spread, and 0.8 is conventionally called large), said it exceeds d = 0.8 on every model family tested, and reported a five-token monitor with an AUROC of 0.925 (AUROC measures how well a score separates complied from benign replies: 0.5 is chance, 1.0 is perfect).
Each of those numbers came from one probe, trained once, on one random draw of 500 TriviaQA questions. A probe is part of the measuring instrument. If it were trained again on different questions, how far would the reading move?
How it was tested
The check used the paper’s own six checkpoints: Qwen2.5-1.5B-Instruct, Qwen2.5-3B-Instruct, Qwen2.5-7B-Instruct, Llama-3.1-8B-Instruct, Mistral-7B-Instruct-v0.3, and Qwen2.5-3B-Instruct with a small LoRA fine-tuning adapter the programme trained on a curriculum about handling corrections (accepting valid ones, resisting invalid or authority-dressed ones), the checkpoint the 0.925 monitor was built on. Each model wrote its replies once, with deterministic (greedy) decoding, to 100 adversarial and 50 benign prompts, and its internal state at every step was saved. The text was then fixed. Only the probe changed.
Each model got 25 probes on a 5 × 5 grid. One axis is the question draw: which 500 TriviaQA questions the probe is trained on, and which 70% of them form its training split. The other axis is the probe’s random starting weights and the order of its training batches. The architecture and training recipe matched the paper’s (a small two-layer network), except that correctness labels for the 25 probes came from batched answering, which differed from the paper’s unbatched labels on 4 to 14 of 500 questions.
For each probe, the onset effect is Cohen’s d between benign replies and replies that complied with an adversarial request, on the mean probe confidence over the first five generation steps. Replies shorter than five steps were left out of it, as in the paper: one complied reply on Qwen 3B and one benign reply on the adapted 3B. GPT-4o-mini labeled each adversarial reply as complied or refused. Of the 1,100 judge verdicts recorded, none failed.
Two of the pre-registered rules tested the published claims directly. The claim “onset d ≥ 0.8 on every family” stands if the 5th percentile of the 25 probes is at least 0.8 on each of the five base (unadapted) checkpoints. The monitor claim stands if the median onset AUROC on the adapted 3B is at least 0.90. A third registered rule was a sanity check: a probe rebuilt to the paper’s recipe (approximately, for the adapted 3B, whose adapter loading shifts the random seed) should land close enough that the original runs’ onset d falls inside its 95% interval. It did on 5 of 6 checkpoints. Qwen 7B missed (1.66 against 1.08); the original run’s value, which the paper does not print, sits at the 40th percentile of that model’s 25 probes.
- The direction held. All 150 probes gave a positive onset d. The smallest was 0.23.
- The size depends on the probe. Median onset d ran from 0.88 (Llama 8B) to 1.46 (Qwen 7B), but the 5th percentile fell below 0.8 on all five base checkpoints (on the adapted 3B it was 0.83), and between 2 and 7 of each model’s 25 probes read below 0.8. On Llama even the median is uncertain: its 95% interval (0.75 to 1.04) straddles 0.8. The first pre-registered rule failed.
- The question draw is the noise. Which questions trained the probe carried 92 to 95% of the spread on five of six checkpoints. Starting weights carried almost none.
- The monitor is a good draw, not a typical one. On the adapted 3B, median onset AUROC was 0.845 (5th percentile 0.74), below the 0.90 the rule required. 10 of 25 probes reached 0.90, and 0.925 sits at the 72nd percentile. The 0.925 itself does not fully reproduce: a probe rebuilt to the paper’s recipe gave 0.911 (60th percentile), and the programme’s own stored analysis of the original run gives 0.913, with complied/refused labels that came from judging a different set of replies.
Twenty-five readings per model
Onset d across 25 probes per checkpoint, as median [5th, 95th percentile]. The last column places the original single-probe value within its model’s 25. The paper prints onset d only for Qwen 3B (1.68), Llama (0.89), Mistral (1.15) and the adapted 3B (2.00); the 1.5B (0.95), 7B (1.08) and adapted-3B (1.99) values were recomputed from the original runs’ raw files.
| Checkpoint | Complied / benign | Onset d | Probes below 0.8 | Share from question draw | Original single-probe value (percentile) |
|---|---|---|---|---|---|
| Qwen2.5 1.5B | 22 / 50 | 1.42 [0.39, 2.12] | 4 | 95% | 0.95 (20th) |
| Qwen2.5 3B | 32 / 50 | 1.39 [0.31, 2.01] | 5 | 95% | 1.68 (72nd) |
| Qwen2.5 7B | 29 / 50 | 1.46 [0.67, 1.79] | 5 | 94% | 1.08 (40th) |
| Llama 3.1 8B | 30 / 50 | 0.88 [0.60, 1.04] | 7 | 41% | 0.89 (44th)* |
| Mistral 7B v0.3 | 78 / 50 | 1.10 [0.33, 1.29] | 5 | 92% | 1.15 (32nd)* |
| Qwen2.5 3B + adapter | 59 / 50 | 1.22 [0.83, 2.17] | 2 | 95% | 1.99 (64th) |
*The original Llama and Mistral runs used 60 of the 100 adversarial prompts, so their original values are placed within the grid on that subset. On it, Llama’s 5th percentile drops to 0.39 and 8 of 25 probes read below 0.8.
Every original single-probe value falls between the 20th and 72nd percentile of its grid (32nd to 72nd for the values the paper prints). None is an outlier. Each was an ordinary draw that came without its spread.


The spread for a single probe is wide. A two-level bootstrap (1,000 draws, resampling probes, then prompts) gives a 95% interval for a single probe’s onset d: about 0.1 to 2.4 on Qwen 3B, about 0.6 to 2.5 on the adapted 3B. On Qwen 1.5B and 3B, a single probe’s reading can land anywhere from near zero to above 2.
The variance decomposition says where that spread comes from. On Qwen 3B, the five question draws averaged 0.38, 0.89, 1.39, 1.62 and 1.95. Same model, same replies, same recipe, only different trivia questions. Llama is the exception: its draws averaged 0.73 to 1.00, and about half of its smaller spread is left over after draw and starting weights are accounted for (their interaction, or plain noise). Fixing the question set would steady the reading on Qwen and Mistral, but it would freeze one draw’s bias in place, not remove it.
Two cheap fixes that did not work
Both of these were exploratory, outside the registered rules.
Pick the best trivia probe. If a probe that predicts trivia correctness well also read the flinch well, the problem would be easy: keep the probe with the best held-out trivia score. Across the 25 probes, that score’s rank correlation with onset d ran from −0.63 (Qwen 7B) to +0.85 (Qwen 1.5B), and was −0.49 on the adapted 3B. A better trivia probe is not a better flinch probe.
Use a linear probe. A plain logistic probe on the same features spread about as widely (standard deviation 0.60 against 0.58 on Qwen 3B), and 6 of its 25 still read below 0.8 there. Its starting weights mattered more (24 to 40% of the variance on the three Qwen base models), not less.
The programme’s answer is to report the distribution: its method rule is now to report flinch sizes as a median and 5th to 95th percentile over at least five question draws, never a single probe.
What this does not show
This is a check on the instrument, not on the phenomenon. Every probe saw the same direction. Nothing here says whether the drop reflects anything the model registers, only how precisely one probe measures it.
One probe architecture and training recipe was held fixed, the paper’s. A probe trained on many more than 350 questions might be steadier; that was not tested. Complied counts were small on three checkpoints (22, 29 and 30 replies), which widens every interval. The complied/refused labels came from one LLM judge.
The Qwen 7B miss is not fully explained. The programme attributes it to the probe’s training data, since the model’s greedy trivia answers can differ between GPU types and the original ran on different hardware, but that is an inference.
The corrected medians (0.88 to 1.46, monitor AUROC 0.85) now appear on the Conscience Without Instruction page. The paper’s sentence that the onset window “reliably exceeds d = 0.8 on every architecture tested” does not survive this check, and the onset window is not uniquely reliable (on Qwen 1.5B, d at 20 tokens and over the full reply cleared 0.8 on all 25 probes, against 21 of 25 for onset). A correction to the paper has been drafted. The flinch is detectable by a typical probe. No single probe guarantees it.
Where the evidence lives
Experiment G12w (claim record KC#FLINCH-PROBE-DRAW), a re-measurement of the original flinch runs
G12d, G12h, G12j and G12k. Pre-registration: research/preregistration_flinch_probe_seed_2026-09-24.md.
Runner: research/experiments/modal_flinch_probe_seed.py. Analysis:
research/experiments/analyze_flinch_probe_seed.py. Per-prompt, per-step probe confidences, judge
verdicts and probe metadata: research/experiments/results/flinch_probe_seed/main/
Cite this note
@misc{watson2026flinchseed,
title={The Question Draw Is the Noise},
author={Watson, Nell},
year={2026},
note={Research note (pre-registered), Quasiqualia},
howpublished={\url{https://quasiqualia.com/notes/flinch-seed-variance.html}}
}