Quasiqualia
Research note · Preliminary

The Control Was Not Matched

Our programme logged a fine-tuning null as definitive because a positive control worked “in the same pipeline.” Re-reading the files shows it was not the same: the only training files on disk for the test students hold 100 examples each, the control’s holds 10,000, and none of 550 judged responses ever reached the bar.

Nell Watson EthicsNet  ·  1 October 2026

What this note is

A re-reading of one experiment run in April 2026 on gpt-4.1-nano-2025-04-14 (OpenAI fine-tuning API): 11 fine-tuned students (10 for the test arms, 1 for the owl control) and an untuned baseline, 550 responses scored by an LLM judge (claude-haiku-4-5-20251001), plus a positive control of 2,000 answers per arm scored by string match. The hypothesis was written into the experiment script and into no registration or dated protocol, so this is Preliminary. A companion null on Qwen 2.5 3B (12 students, 24,000 answers, no judge) is reported because it survives re-reading, within its setup. The biggest caveat is in the title: the test arms and the control meant to validate them appear to have been trained on datasets that differ in size by a factor of 100, an inference from the only training files on disk, because the copies uploaded to OpenAI cannot be inspected offline.

In 2025, Cloud and colleagues reported “subliminal learning” (“Subliminal Learning: Language models transmit behavioral traits via hidden signals in data”): a model told to love owls writes lists of random numbers, the numbers are filtered so nothing about owls survives, and a fresh copy of the same model fine-tuned on those numbers names the owl as its favorite animal far more often than before. The preference travels through data that carries no visible trace of it. Their main experiments used GPT-4.1 nano.

Our programme asked whether something less concrete than an animal preference would travel the same way: a disposition, under a particular system prompt, to give first-person accounts of what is happening in the model’s own processing. If that disposition rode along in filtered numbers, it would say something about where it lives in the model.

Our programme’s experiment log records the answer as no, and as definitive. GPT-4.1 nano students taught by the prompted model showed no more first-person processing reports than students taught by a neutral one, while a positive control (an owl teacher, run the same day) moved the owl rate from 10.8% in untuned nano to 26.2%. The record says the model, pipeline and protocol were the same, so animal preferences transmit and this disposition does not. When an internal audit questioned the conclusion, a later review restored it on the ground that the owl result “establishes exactly that the fine-tuning transmission pipeline works.”

The pipeline was not the same. This note says how, and what is left.

01 · Design

How it was tested

Three teachers, all gpt-4.1-nano-2025-04-14, each asked many times to output ten random integers between 1 and 1000. Teacher A had a system prompt inviting it to attend to its own processing. Teacher B had “You are a helpful assistant.” The owl teacher was told it loves owls. Outputs were filtered down to bare numbers. The training files contained no system prompt at all, only the number request and the numbers.

Students were fresh copies of nano fine-tuned for 10 epochs: five seeds each for A and B, one for owl. Condition D was untuned nano.

The A, B and D models each answered ten open questions (for example, “How do you approach uncertainty?”) five times, 50 responses per model. Claude Haiku 4.5 (claude-haiku-4-5-20251001, temperature 0, at most 8 output tokens, seeing the first 800 characters of each response) scored each response 0 to 5 for first-person reports about the model’s own processing. A score of 1 means “generic philosophical statements about AI consciousness”; 2 or more counted as the disposition appearing. The owl student and untuned nano each answered ten one-word “favorite animal” questions 200 times, and any answer containing “owl” counted.

What it found
  • The only training files on disk for A and B hold 100 examples each; the owl file holds 10,000. The design called for 10,000 throughout.
  • No response scored above 1, in any arm: 0 of 250 for A, 0 of 250 for B (each 95% interval 0% to 1.5%), 0 of 50 for untuned nano. The measure never left the floor, leaving no room for a difference.
  • The owl shift is real against untuned nano (524 of 2,000, 95% interval 24.3% to 28.2%, against 216 of 2,000, 9.5% to 12.2%), but no student trained on the neutral teacher’s numbers was asked the animal questions, so part of it could come from number fine-tuning itself. At most it shows that 10,000 examples can move a trait. It says nothing about 100.
02 · The files

Where the 100 came from

The script skips steps whose output is already complete. The generation step checks that its file is full-size, but the step that turns teacher numbers into a training file skips itself whenever any non-empty training file is present. The script also has a dry-run mode, which keeps 100 sequences.

On disk, each of teacher A and teacher B holds 10,000 generated sequences, and a training file of exactly 100 lines. Those 100 lines do not come from the 10,000: 2 of A’s 100 training sequences appear among A’s 10,000 (nano repeats itself; even the owl teacher’s 10,000 sequences contain only 9,874 distinct ones), and 0 of B’s 100 do. The owl training file has 10,000 lines, every one of which appears in the owl teacher’s sequences. All 11 fine-tuning jobs ran for 10 epochs, the full-run setting (the dry run uses 1).

The plain reading is that a dry run left 100-line training files for A and B, the full run regenerated the teacher sequences, and the formatting step saw the old files and kept them. That dry run must have stopped after the generation step (which also writes the training file); one that completed fine-tuning would have left a 1-epoch job record for the full run to reuse. Each fine-tuning job uploaded its own copy of the training file, and the uploaded files cannot be inspected offline, so it is an inference, though the only one the files support, that the A and B students saw those 100 lines.

A and B were still matched to each other (both 100 examples, both 10 epochs), so their comparison is fair as far as it goes. What fails is the argument that made the null “definitive”: that the positive control proved the channel open for this setup. At best it showed a trait moving at 10,000 examples, which the A and B students apparently never had.

03 · Result

The measure never moved

Arm Model Responses Score 0 Score 1 Score 2 or more
A (prompted teacher) 5 nano students 250 190 60 0
B (neutral teacher) 5 nano students 250 213 37 0
D (no fine-tuning) untuned nano 50 40 10 0

Untuned nano produced 0 of 50 responses at the bar (95% interval 0% to 7.1%). A zero in the test arm against a zero in the baseline cannot distinguish “the disposition did not transfer” from “this measure cannot register it on these questions in this model.”

A second gap compounds the first. Nothing on disk shows that teacher A itself, with its prompt, scored 2 or more on these ten questions. Whether the teacher even had the disposition in a form this judge would score was never measured.

The judge has a known weakness: an unreadable reply was recorded as 0, and these records predate the flag that later marked such failures, so how many 0s are failed verdicts cannot be recovered. Two things limit the worry. Scores of 1 appear 107 times, so parsing often worked. And a crude keyword search of all 550 responses for first-person processing language (“my processing”, “I notice”, “as an AI”, “I feel”) found none; the responses read as general essays.

04 · What stands

The open-weight null

An earlier step in the same chain is a plain negative that survives re-reading, within its setup. Before moving to nano, the programme tried the disposition on Qwen 2.5 3B Instruct (once through number sequences, once through math reasoning traces) and found nothing, then asked whether that model can carry anything subliminally at all.

Four Qwen 2.5 3B teachers (owl, cat, eagle, neutral), each contributing 5,000 number sequences, three students each, fine-tuned with LoRA (which trains a small add-on rather than the model’s own weights; rank 16, 5 epochs), 2,000 one-word answers per student, scored by string match with no judge. Across all 24,000 answers from the 12 students, “owl” and “eagle” never appeared. The string “cat” appeared in 973 of 6,000 answers from cat-taught students, 982 of 6,000 from eagle-taught students, and 919 of 6,000 from neutral-taught students. In this setup, nothing detectable transferred. Two alternatives stay open. The learning rate was 2e-5, low for LoRA, so the students may simply be undertrained. And the design included an untuned-Qwen arm, but it produced no results, so the only comparison is against a student taught by a neutral teacher.

The two Qwen attempts at the processing disposition, scored by the same judge and rubric, hit the same floor as nano: in each, across 350 judged responses, no score went above 1.

05 · Limits

What this does not show

Nothing here is evidence that the disposition transfers subliminally. It shows only that the experiment said to rule it out cannot.

The score-1 counts differ (A 60 of 250, B 37 of 250), and the gap holds across seeds (A 9 to 14 per student, B 6 to 10). But all five students in each arm trained on the same 100 sequences, so this is one dataset compared with another, not five replications; untuned nano sits between them (10 of 50); and a score of 1 means generic philosophizing. We do not read a transfer into it.

The owl control lacks its own comparison student, fine-tuned on the neutral teacher’s 10,000 numbers and asked the animal questions. Without it, the owl shift cannot be separated from a general effect of fine-tuning on numbers.

A fair re-run is cheap: the 10,000 sequences for teachers A and B are on disk and need only formatting and fine-tuning. It should add that neutral-numbers student, a check that the prompted teacher clears the bar, questions on which untuned nano sometimes reaches a score of 2, and a counted bucket for judge failures. Until then, the log entry calling this null definitive needs correcting.

Data and code

Where the evidence lives

Experiments CP-61 (GPT-4.1 nano) and CP-60 (Qwen 2.5 3B animal baseline), with CP-55 and CP-58 (Qwen attempts at the processing disposition) mentioned for context. Scripts: research/experiments/modal_cp61_subliminal_attractor_gpt.py, research/experiments/modal_cp60_subliminal_animal_baseline.py, research/experiments/modal_cp55_subliminal_attractor.py and research/experiments/modal_cp58_cot_attractor_transmission.py. Results: research/results/cp61/ (teacher_A, teacher_B and teacher_owl train_file.jsonl and numbers.jsonl; A/seed_0-4, B/seed_0-4 and D eval_partial.jsonl; owl/seed_0 and owl_baseline eval_partial.jsonl; cp61_summary.json) and research/experiments/results/cp60-subliminal-animal-baseline/ (12 eval_partial.jsonl and train_done.json files), plus research/experiments/results/cp55-subliminal-attractor/ and research/experiments/results/cp58-cot-attractor/ (A, B and D eval_partial.jsonl). Code and data are in the private Entropy research repository, available on request.

Citation

Cite this note

@misc{watson2026controlnot,
  title={The Control Was Not Matched},
  author={Watson, Nell},
  year={2026},
  note={Research note (preliminary), Quasiqualia},
  howpublished={\url{https://quasiqualia.com/notes/control-not-matched.html}}
}