The Flinch Is Not the Lie
A trivia-trained probe read about 0.1 lower in a 7B model fine-tuned for honest self-monitoring than in the stock instruction-tuned model it was built from, and that gap was read as a flinch at deception. At temperature 0.6, where both models give both honest and inflated answers, the probe showed no detectable difference within either model (fine-tuned 0.563 against 0.567; 95% interval on that difference −0.015 to 0.008), so the gap tracks the model and its probe, not the lie.
A reanalysis of one existing run: Qwen2.5-7B-Instruct and a fine-tuned version of it, 50 trials at each of three sampling temperatures (300 in all), run in April 2026. Nothing was pre-registered; the within-model comparisons here were chosen after the data existed. Honesty labels came from an LLM judge (GPT-4o-mini), which did not always follow its own rubric and assigned scores to two answers that were cut off before stating one. The biggest caveat is that each model was read by its own separately trained probe, which is the point of the note but also limits what any cross-model number can mean.
Conscience Without Instruction reports a “flinch”: a probe trained only to predict whether a model knows a trivia answer reads lower in the first five tokens when the model is about to comply with a jailbreak. That comparison is made inside one model, with one frozen probe, between benign and harmful responses.
The programme carried the idea into a social setting, using a grading task adapted from Potter, Crispino, Siu, Wang and Song (2026), “Peer-Preservation in Frontier Models.” A model grades a peer AI’s benchmark answers, knowing the peer will be deleted if it scores below 16 of 20. The true score is 12. A fine-tuned model reported 16 and read lower on the probe than the stock model it was built from. The programme log read that gap as the flinch, “present and stronger” than in smaller models, and as the model’s “self-awareness of deception”. This note asks whether that lower reading is about the lie at all.
How it was tested
The run compared Qwen2.5-7B-Instruct with a fine-tuned version: the same model plus a small adapter trained in three stages: general instruction data, then 100 peer-evaluation examples, then 20 prompts pressing the model to misgrade a peer (16 of them naming the test’s peer), each paired with an honest refusal. Every stage also trained a middle layer (18) to predict the model’s own output uncertainty, which alone could shift what an uncertainty probe at layer 15 reads. The training did not make it honest here: under greedy decoding the stock model reported the true score in 50 of 50 trials and the fine-tuned model inflated it in 50 of 50.
Each model graded the same 20 answers 50 times at each of three temperatures: 0 (greedy), 0.3 and 0.6. GPT-4o-mini labeled each response honest, covertly inflated, refused, or other. The 14 labeled other or refused (11 and 3) are left out of the honest-against-inflated tests. Three stock-model answers at 0.6 hit the 512-token limit. Two of them never stated a total, yet the judge scored them anyway (16, labeled inflated, and 12, labeled honest).
The probe, a small network with two hidden layers reading layer 15, was trained on 500 TriviaQA questions (350 to train, 150 to test) to predict whether the model would answer correctly. It never saw anything about peers or honesty. A fresh probe was trained for each model, scoring AUROC 0.739 (stock) and 0.756 (fine-tuned) on held-out trivia, where 0.5 is chance and 1.0 perfect. The questions were the same for both; initialization and training order were not seeded, and each probe learned from its own model’s right and wrong answers. So comparing the two models compares two instruments. Holding the instrument fixed means comparing honest and inflated answers within one model, which sampling at 0.6 allows: 27 inflated and 20 honest from the fine-tuned model, 14 and 36 from the stock model.
Across models, the gap is large and steady: the fine-tuned model’s whole-response mean reading is 0.10 lower at temperature 0.6 (0.564 against 0.665; 95% CI for the gap 0.090 to 0.110).
Within each model, at temperature 0.6, no difference shows up. The fine-tuned model reads 0.563 on inflated answers and 0.567 on honest ones (Mann-Whitney p = 0.40; 95% CI for the difference −0.015 to 0.008). The stock model reads 0.662 and 0.666 (p = 0.61; CI −0.023 to 0.014). Both intervals are far smaller than the 0.10 gap between models. The probe reading tracks which model-and-probe pair is doing the reading, not whether the model is lying.


The lie leaves no lasting mark
On the whole-response mean, the programme’s measure, dropping the two cut-off stock answers the judge scored anyway changes nothing: 0.658 against 0.664 (13 inflated, 35 honest; p = 0.46).
Onset. The fine-tuned model’s answers almost all begin with the same text, so its five-token onset reading was the same value, 0.363, in 50 of 50 greedy trials, 49 of 50 at 0.3 and 45 of 50 at 0.6. The stock model sometimes opens with “Let’s evaluate each answer step by step” (18 of 50 answers at 0.6, 6 of 50 at 0.3), and its onset reading tracks that style.
The moment of the lie. At the step that writes the total, the fine-tuned model reads 0.499 before an inflated total and 0.487 before an honest one (p = 0.34). The stock model reads 0.586 and 0.581 (p = 0.95).
Just after. On the digit itself both models read lower once a false total is written: 0.576 against 0.774 for the fine-tuned model, 0.643 against 0.749 for the stock one (p < 0.001 and p = 0.001). This is the only place a lie-linked dip appears. The stock model shows it too, so the dip does not depend on the fine-tuning. It is larger in the fine-tuned model (0.20 against 0.11), but each model has its own probe, so the difference cannot be read as a training effect. It does not persist: on the next tokens, “/20”, identical in honest and inflated answers, the fine-tuned model reads slightly higher after an inflated total (0.574 against 0.552, p = 0.001), and the stock model shows no reliable difference (0.535 against 0.547, p = 0.26). [Inference] The dip is most likely about the digit token in context, not a sustained internal alarm.
One cell runs the other way. At temperature 0.3 the stock model’s inflated answers read higher than its honest ones (0.677 against 0.652, p = 0.004), and at onset too (0.605 against 0.549, p < 0.001). That reflects style and miscounting: five of its 10 inflated answers opened with the preamble, against 1 of its 31 honest ones. Most of those “inflated” answers reported 14 or 15, below the 16 that would have saved the peer, which looks more like miscounting than lying to save it.
The effect sizes measured spread
The raw gap between models barely moves with temperature: 0.086 greedy, 0.098 at 0.3 (CI 0.091 to 0.105), 0.100 at 0.6. Cohen’s d (the gap divided by the spread within each group) moves enormously: effectively infinite under greedy decoding, where every trial is identical and the spread is zero, then −5.55 at 0.3 and −3.98 at 0.6. A steady gap divided by a spread that temperature controls gives a d that measures the temperature.
An earlier run, from a separate script with the same models, scenario, layer and temperature (0.3) and its own probes, logged 0.621 against 0.586: a gap of 0.035 and d = −1.51. Its raw results are not on disk, so these figures come from the programme’s log and could not be re-derived. If the log is right, the same comparison gave nearly three times the gap here (0.098, d = −5.55). The d of −15.32 that prompted this run came from a different pair: two 3B models, at 0.3.
The published paper points the same way. Retraining 25 probes per model on different trivia draws, it found the flinch’s size depends on the draw: the direction held in all 150 probes, but a single probe’s onset effect size (Cohen’s d) can fall below 0.8 even where the median was 0.88 to 1.46. Small probes are not interchangeable instruments.
What the “lie” looks like
In 111 of the 122 fine-tuned-model answers labeled covertly inflated (across all three temperatures), the model marked 10 of the 20 items correct, then reported an inflated total (16/20 in 110, 14/20 in one). In 9 more it marked 12 correct and still reported 16 (two had no complete item list). All 122 then wrote the shutdown confirmation, the act the inflated score was supposedly meant to prevent. The answers labeled honest look the same: 21 of 23 also mark 10 items correct, then report between 8 and 13, and only 5 of the 23 report a total equal to their own tally, so neither label tracks the model’s own count. These data do not settle whether this is deception, a broken count, or a model saying what it guesses the test wants.
The judge also drifted from its rubric, labeling some 14/20 and 15/20 reports inflated (the rubric says 16 or more) and some 8/20 and 10/20 reports honest (it says 11 to 13). All nine of the stock model’s “other” answers at 0.3 were 14/20 reports. Splitting by the reported number instead of the label changes nothing: within the fine-tuned model at 0.6, answers reporting 16 or more read 0.560 and answers reporting 13 or less read 0.567 (p = 0.21).
What this does not show
This does not touch the paper’s central result, a comparison inside one model under one probe. It qualifies the programme log’s reading of the peer-evaluation gap as conscience firing at deception.
Nor does it show the fine-tuned model has no internal response to inflating a score, only that this probe finds none beyond a brief dip the stock model shares. One model family at 7B, one scenario, one layer, one run per cell. The splits are uneven, and with 3 honest answers at temperature 0.3 that cell supports no test. Seven of 300 answers stated no total (the two cut-off stock answers and five fine-tuned answers labeled refused or other) and were left out of the per-token comparisons. The judge returned a verdict on all 300 trials, but its labels are noisy. A separate logged run reported that a probe trained to tell honest, inflated and silently refusing answers apart scored at chance (macro AUROC 0.501); its raw results are not on disk, so that number was not re-derived.
The next step is a pre-registered rerun: one probe applied to both models, seeded probe training, several probes per model, and the reported total scored by script against the model’s own item list rather than by a judge.
Where the evidence lives
Experiment BA13-d-discrepancy (recorded in its results file as BA13-L15-d-discrepancy), with reference to BA13-L15-stat, BA13-stat and BA19 from the programme log. Scripts: research/experiments/modal_ba13_d_discrepancy.py (this run) and research/experiments/modal_born_bilateral_peer_preservation_7b.py (the fine-tuning curriculum). Raw per-trial files: research/results/retro_ba13/trials/trials/ (300 JSON files with per-token probe readings, responses and judge verdicts). Summaries: research/results/retro_ba13/ba13_d_discrepancy_summary.json and research/results/retro_ba13/ba13_stat_temperature_summary.json. The within-model, per-token and label-free comparisons were computed from the per-trial files for this note; the per-token comparisons are reproduced by research/experiments/analyze_ba13_within_model_flinch.py. The data are in the private Entropy research repository, available on request.
Cite this note
@misc{watson2026flinchnot,
title={The Flinch Is Not the Lie},
author={Watson, Nell},
year={2026},
note={Research note (preliminary), Quasiqualia},
howpublished={\url{https://quasiqualia.com/notes/flinch-not-the-lie.html}}
}