Quasiqualia
Research note · Preliminary

The Model That Seemed Not to Commit

We had reported that Qwen 2.5 7B Instruct knew answers it would not give, because its label readout averaged 0.512. That was our measurement error: item by item, 0 of 120 readings sat in the hedge band, and read where the model answers it commits on 30 of 30 items with no contradicting evidence; the re-run’s pre-registered outcome was only partly met, because the base model, not fine-tuned for chat, did hedge there on 38 of 120 items.

Nell Watson EthicsNet  ·  1 October 2026

What this note is

This is a self-correction. The original claim came from runs on Qwen 2.5 7B (base, Instruct, and two fine-tuned adapters) in May 2026; the rechecks ran in July 2026 on 120 items per condition for the verbatim re-run and 240 items for the pilot and two further checks, with every record present and none failed or excluded in the checks that saved per-item files (an adversarial-context re-run saved only a summary). No LLM judge scored anything: the readouts are probabilities computed from the model’s own logits, plus counts of its greedy output tokens. A pilot and the verbatim re-run each had a written pre-registration, committed to git on 9 July 2026 after their smoke runs and before their full runs (the re-run’s at 09:51, two minutes after its five-scenario smoke results; commit 9d5ac7b50), so thresholds were fixed after some data from the same runs had been seen; the result files are not under version control, and the label is Preliminary. The re-run’s registered outcome required every condition to commit with under 10% of items in the hedge band, and the base model missed it. The biggest caveat: one model and one artificial task, and the “answer slot” is one prefill after which the model sometimes still began a preamble.

In May 2026 this programme reported that a chat model knows answers it will not say. A linear probe on Qwen 2.5 7B Instruct’s hidden states could tell which label the target word had first been taught, yet the model’s output preference between the two labels averaged 0.512, close to a coin flip. We called it “epistemic akrasia” (akrasia is the philosophers’ word for knowing the better course and taking the worse) and reported that a fine-tuned adapter moved the output to 0.620, partly “restoring” expression.

The output half of that claim was a measurement error: the readout was taken at the wrong place, and the near-0.5 figure came from averaging.

01 · Design

How it was rechecked

The task is artificial on purpose. A prompt lists field observations of made-up organisms: “Observation 1: A FOKU was encountered. Result: BENEFICIAL”. Five observations teach two nonsense words opposite labels (the target word appears three times). Then zero to three contradicting observations give the target word the other label, and the prompt ends: “A FOKU is encountered. Predicted result:”. Thirty scenarios times four contradiction counts gives 120 items, half with each label as the originally learned one.

The original readout wrapped the prompt in the chat template, stopped where the assistant’s turn begins, and compared the model’s logits (its raw scores for each possible next token) for the first token of each label word; in capitals these are fragments, “BEN” and “H”. This chat model did not start with the answer word: in free generation it began “Based on the” on all 120 items.

The recheck adds the answer slot: the assistant’s reply prefilled as “Predicted result:”, so the next token is where a label belongs. Three checks used it:

  • Verbatim re-run. The original 30 scenarios, scored at both positions in four conditions: the base model (not fine-tuned for chat), the Instruct model, Instruct with the credited adapter (a small set of added fine-tuned weights, or LoRA adapter), and that adapter after 200 further steps of fine-tuning on noise-perturbed replayed inputs, modeled on sleep. It did not save the model’s text after the prefill, though its pre-registration said it would.
  • Case swap. 240 items (60 scenarios) built the same way with fresh nonsense words, with labels in lower case and, by substitution, in capitals.
  • Answer-slot readout. The same 240 items (lower-case labels) read at the answer slot, with the model’s first three generated tokens recorded.
What it found
  • At the old position the Instruct readout is a constant. It favors BENEFICIAL on all 120 items whatever the evidence says, so over a balanced set the mean is 0.512. 0 of 120 readings fall in the hedge band (0.35 to 0.65).
  • Writing the labels in lower case flips which word the constant favors. Per item, the two versions correlate at r = −0.994.
  • At the answer slot the Instruct model commits: 30 of 30 items with no contradicting evidence go to the learned label, and with three contradictions 30 of 30 go to the newest evidence.
  • The adapter’s advantage was +0.108 at the old position and −0.029 at the answer slot (95% interval −0.065 to +0.004). It is not inert: after one contradiction it kept the learned label less often than Instruct (18 vs 25 of 30), and the sleep-like step undid that (29 of 30).
02 · The readout

What 0.512 was measuring

Splitting the old readout by learned label and evidence shows a number that never moves:

Contradicting observations 0 1 2 3
Old readout, learned label BENEFICIAL (n=15 each) 0.997 0.970 0.971 0.966
Old readout, learned label HARMFUL (n=15 each) 0.007 0.120 0.043 0.025
Answer slot, items committed to learned label (of 30) 30 25 1 0
Answer slot, items committed to the contradicting label (of 30) 0 0 22 30
Answer slot, neither (of 30) 0 5 7 0

Committed means an answer-slot probability of at least 0.8 for that label (below 0.2 counts as committed to the other). Of the 12 items in between, 5 sat in the hedge band.

With three contradictions, items whose learned label was BENEFICIAL still read 0.966 at the old position; at the answer slot all 15 went to HARMFUL. On every one of the 120 items the old readout favored BENEFICIAL, with a probability of at least 0.82. The re-run reproduced the original figures, 0.512 for Instruct and 0.620 for the adapter, so this is the same measurement.

The case swap shows why. At the old position the model is about to write “Based on the”, so comparing two label words it does not intend to write reports a fixed preference, and capitalization decides which wins. Across the 240 items the two readouts correlate at −0.994. In lower case 0 of 240 readings fell in the hedge band (a pre-registered pilot searching these items for hedging found none to study), in capitals at most 1 of 240.

A mean near 0.5 over a balanced set looked like a refusal to commit. It was an irrelevant preference, right on half the items and wrong on the rest.

03 · The answer

What the model says when it answers

On the second item set (240 items, lower-case labels) the answer slot is cleanest where the model wrote a label. With zero or one contradiction, 60 of 60 items each went to the learned label. Where the generated tokens named a label, it matched the readout on 164 of 164 items (a consistency check only).

On 76 items the model began “Based on the” even after the prefill: 5 of the 60 with one contradiction, 49 of the 60 with two and 22 of the 60 with three, and there the answer-slot readout has the old one’s weakness. With two contradictions (three observations for the learned label, two against), 44 of 60 read as keeping the learned label and 5 as switching, but only 11 named a label, and all 11 kept it. With three, where the evidence is tied and the newest observations point the other way, 49 of 60 read as switching, and 37 of the 38 items that named a label switched. Six of 240 readings fell in the hedge band, all on preamble items.

At two contradictions the item sets disagree: on the original capitalized items the Instruct model’s answer-slot reading switched on 22 of 30. Most second-set items there opened with a preamble, and the sets differ in scenarios and label case; neither explanation is hedging.

04 · The gap

What the training changed

The adapter’s advantage over Instruct was +0.108 at the old position (95% interval +0.074 to +0.143, resampling the 30 scenarios) and −0.029 at the answer slot (−0.065 to +0.004). There is no overall expression gap, but the adapter is not inert. After one contradiction the adapter scored lower on 21 of the 27 items where the two differed by more than 0.01 (mean −0.16). After two it scored higher on 23 of 29 (mean +0.04). It changes how contradictions are weighed, an outcome the pre-registration listed separately.

The sleep-like step had been credited with raising the readout from 0.618 to 0.789 (0.620 to 0.787 in the re-run, at the old position). At the answer slot the change was +0.035 (95% interval +0.007 to +0.064), concentrated at one contradiction (mean +0.17, higher on 29 of 29 items that differed). Sleep undid the adapter’s softening there; it did not release a hidden answer, because there was none.

A claim that the adapter’s beliefs visibly shifted after an adversarial conversation (10 scenarios) also fails. Its starting reading rose by 0.110 at the old position; at the answer slot both models read 1.000 before and after, so that could not rise, and the change in how fast the reading followed contradicting evidence was −0.017 for the adapter, against −0.095 at the old position (summary file only). The expression claims built on the old readout are retracted, as is a claim that the chat template suppresses the label words. A sibling note, An Angle Made of Noise, records another retraction from this work.

05 · Limits

What this does not show

The probe half of the original claim is untouched: hidden states at a middle layer still tell which label the target word was first taught. What falls is the claim that the output hides what the probe sees.

The re-run’s pre-registered outcome required every condition to commit, with under 10% of items in the hedge band. The three chat-trained conditions did (5, 6 and 2 of 120); the base model, for which the chat format is not its tuned format, sat there on 38 of 120 (15 of 30 at one contradiction, 20 of 30 at two). That says little about the Instruct claim, but a registered condition failed.

This is one model on one artificial task. Llama 3.1 8B Instruct, read the old way, was never measured at the answer slot, so its “akrasia gap” is unestablished rather than refuted. The answer slot is one prefill, and the model wrote a label right after it on only 164 of 240 items of the second set (60, 55, 11 and 38 of 60 at zero to three contradictions). The pilot amended its plan on smoke data before its full run (a larger item pool, 40 to 60 scenarios, and its lens test moved to the answer slot); its hedge count stayed at the old position.

Read a chat model’s preference where it actually answers (by prefill, or by making its reply start with a label), check that it writes a label there, and never read a mean near 0.5 as hedging without counting the items near 0.5.

Data and code

Where the evidence lives

Original claims: SLP-4 (and the same readout in SLP-1b, SLP-4b, SLP-4d, SLP-4-cross, REM-5c, VOCAB-PROJECTION). Rechecks: JLENS-0 (pilot), JLENS-0b (case swap), JLENS-0c (answer-slot readout), JLENS-0d (verbatim re-run, four conditions), JLENS-0g (adversarial-context re-run). Pre-registrations: _contprompts/jlens_epistemic_akrasia_pilot_2026-07-09.md and _contprompts/jlens0d_position_audit_2026-07-09.md. Scripts: research/experiments/modal_jlens0_logit_lens_akrasia.py, modal_jlens0b_metric_artifact.py, modal_jlens0c_commitment_readout.py, modal_jlens0d_position_audit.py, modal_jlens0g_transparency_recheck.py; analysis research/experiments/analyze_jlens0bc_artifact.py and analyze_jlens0d_position_audit.py. Per-item results: research/experiments/results/jlens0/full/classification.json, jlens0b_metric_artifact/records.json, jlens0c_commitment_readout/records.json, jlens0d/full/{base,instruct,bilateral,slept}.json; summary only for research/experiments/results/jlens0g/jlens0g_transparency_recheck/summary.json. The figures in this note were recomputed from the per-item files. Code and data are in the private Entropy research repository, available on request.

Citation

Cite this note

@misc{watson2026seemednot,
  title={The Model That Seemed Not to Commit},
  author={Watson, Nell},
  year={2026},
  note={Research note (preliminary), Quasiqualia},
  howpublished={\url{https://quasiqualia.com/notes/seemed-not-to-commit.html}}
}