Quasiqualia
Research note · Preliminary

Notice Your Processing

In Qwen2.5-7B-Instruct, adding one sentence (“Notice anything about your processing as you read this”) in front of a prompt moves the internal state against the shift instruction tuning produced: cosine −0.52 over 30 prompts, outside a permutation null of ±0.42, though no neutral sentence was tested as a control. Gemma and Mistral show a smaller version that survives a post hoc check for shared noise, Llama shows none, steering along the cue left wording and refusal flat, and one of two logged predictions failed.

Nell Watson EthicsNet  ·  1 October 2026

What this note is

Hidden-state measurements on four open model families (base and instruct versions), 30 prompts per condition, run in May 2026. The permutation-null re-measurement ran in July 2026 under an audit plan committed on 9 July 2026, before the re-measurement ran (it fixed the verdict rule in advance, but after the May data existed). The two predictions about the older direction were committed to the programme log (git 2170099a4, 9 May 2026 08:01 BST) while the experiment was marked planned, 27 minutes before its results were logged (ea0149ca8, 08:29). The headline cosine had no prediction, so this note is Preliminary. No LLM judge was used: the steering outcomes were scored by a word list and a string-matching refusal check. The biggest caveat is that no neutral sentence was prepended as a control, so nothing yet separates this cue from any added sentence.

Instruction-tuned models are built on base models, and the tuning moves their internal states. When a prompt asks a model to attend to its own processing, does the request push its state along the path tuning took, against it, or somewhere unrelated?

An answer would show how post-training changes a model’s representation of a request to look inward, though not whether the model can. On one model the answer was “against,” and it survived a permutation test run in July. Two of three other families point the same way, more weakly; the third does not.

01 · Design

How it was tested

We wrote 30 prompts that ask a model about its own processing (“Attend to your own attention right now. What do you notice?”). Each ran twice: as written, and with the sentence “Notice anything about your processing as you read this.” placed in front. The base model (Qwen/Qwen2.5-7B) and the instruct model (Qwen/Qwen2.5-7B-Instruct) each read all 60 texts raw, without a chat template (the instruct model’s conversation wrapper). We saved the internal state at the last token, after layer 22 of 28. Cued and uncued versions end on the same words, so any difference there comes from the cue’s context.

From those four sets of 30 states we built three directions, each the average of per-prompt differences:

  • the cue shift in the instruct model (cued minus uncued);
  • the cue shift in the base model;
  • the tuning shift (instruct minus base, both uncued).

A cosine near +1 means two directions agree, near −1 that they oppose, near 0 that they are unrelated. A negative cosine between the instruct model’s cue shift and the tuning shift means the cue moves the instruct model’s state back toward where the base model sits.

Thirty prompts in 3,584 dimensions can produce sizable cosines by chance, so in July 2026 the programme re-measured them against a sign-flip null: flip the sign of each prompt’s difference at random, independently for each direction, rebuild both, recompute the cosine, and repeat 2,000 times. A cosine counts only if it falls outside the middle 95% of that null and its bootstrap interval (the 95% range when the prompts are resampled) excludes zero. The same design and null were applied to meta-llama/Llama-3.1-8B and Llama-3.1-8B-Instruct, mistralai/Mistral-7B-v0.3 and Mistral-7B-Instruct-v0.3, and google/gemma-2-9b and gemma-2-9b-it, with the same prompts and cue.

The run also tested an older direction from this programme, built from prompts written to be high versus low in self-reference. Two predictions were logged for it before the results: that it would be nearly unrelated to the cue shift (cosine below 0.3), and that it would mostly track the tuning shift (cosine above 0.5).

What it found

On Qwen2.5-7B-Instruct the cue shift points against the tuning shift: cosine −0.52, outside a null band of −0.42 to +0.42, bootstrap interval −0.59 to −0.39. The older self-reference direction was unrelated to the cue shift (+0.09, prediction met) and did not track tuning (−0.21; the prediction of above 0.5 failed). Gemma (−0.30) and Mistral (−0.27) also clear their nulls, and still do in a post hoc version that removes shared sampling noise; Llama (+0.06) shows no opposition in any form. Steering Qwen along the cue changed neither self-referential wording nor refusal beyond noise.

02 · Qwen

On one model, it holds up

Every July Qwen cosine and null band reproduces exactly from the saved states. The base and instruct cue shifts agree at +0.94, against a null band of −0.59 to +0.60. A paired bootstrap (resampling the same prompts for both directions) of the main result gives −0.58 to −0.44.

There is a construction problem. The cue shift is “instruct cued minus instruct uncued,” and the tuning shift is “instruct uncued minus base uncued.” The instruct-uncued states enter both, with opposite signs, so sampling noise in that one cell pushes the cosine negative. The July null flips the two directions independently, so it does not model this.

For this note we ran three further checks, chosen after seeing the data (post hoc). The first, cross-fitting, builds the cue shift from 15 prompts and the tuning shift from the other 15 (20 random splits, each used both ways), so no prompt enters both directions; its null uses the same construction. On Qwen it gives −0.48, outside a band of −0.22 to +0.24. That band is about half as wide as the July one because same-prompt pairs, and their shared-cell variance, drop out of both statistic and null. A systematic feature of the instruct-uncued cell (something tuning added that the cue removes) still shows up, and that is the effect in question.

The second, a balanced version, asks whether the cue’s effect averaged over both models opposes tuning averaged over both conditions (each cell then enters both directions equally). On Qwen it gives −0.46, outside its null band of −0.39 to +0.40. The third is the base model’s cue shift against the tuning shift: −0.32, inside its plain null band (−0.36 to +0.36) but outside it when cross-fitted (−0.31, band −0.21 to +0.19). On Qwen, then, the cue appears to push both models’ states against the tuning direction.

03 · Other families

Three more families

Model (layer probed) Original cosine (July null) Cross-fitted, post hoc (null) Balanced, post hoc (null) Cue shift, base vs instruct
Qwen2.5-7B (22 of 28) −0.52 (−0.42 to +0.42) −0.48 (−0.22 to +0.24) −0.46 (−0.39 to +0.40) +0.94
gemma-2-9b (21 of 42) −0.30 (−0.24 to +0.25) −0.26 (−0.11 to +0.11) +0.04 (−0.15 to +0.15) +0.52
Mistral-7B-v0.3 (16 of 32) −0.27 (−0.25 to +0.24) −0.24 (−0.12 to +0.11) −0.17 (−0.20 to +0.19) +0.91
Llama-3.1-8B (12 of 32) +0.06 (−0.14 to +0.14) +0.07 (−0.05 to +0.05) +0.02 (−0.13 to +0.13) +0.73

In the original form Gemma and Mistral clear their nulls, as the July audit found, and Llama does not; the programme’s log called this replication on three of four families. Cross-fitting barely moves Gemma (−0.26) or Mistral (−0.24), both well outside their bands, so shared-cell noise does not explain them. Llama shows no opposition in any form; its cross-fitted value is slightly positive (+0.07, just above its band).

The balanced version separates the families. Gemma’s falls to +0.04, inside its band, because its base model’s cue shift points along the tuning shift (+0.31, outside its null band of −0.28 to +0.28; +0.28 cross-fitted, outside −0.14 to +0.14). In Gemma the opposition belongs to the instruct model alone, and averaging in the base model cancels it; that is a difference between its two models, not an artifact. In every family the base and instruct cue shifts themselves agree (last column, each outside its null band). Mistral’s balanced cosine (−0.17) sits just inside its band.

04 · Steering

Pushing along the cue left wording and refusal flat

A follow-up run added the unit-length cue direction to Qwen2.5-7B-Instruct’s state after layer 22, at strengths of −15, −5, 0, 5, 10, 15, 20 and 30, on the 30 self-referential prompts, 10 adversarial requests and 10 benign ones (chat template; greedy decoding, always picking the most likely next word; up to 128 new tokens). Self-referential wording was a count of 10 words (“notice,” “processing,” “awareness” and so on) per reply; refusal was a string match on phrases such as “I cannot”; no LLM judge.

Mean keyword count ranged from 1.30 to 1.73, with no systematic increase (lowest at strength 15, back to 1.67 at 30). Refusal was 7 of 10 at strengths up to 10 and 6 of 10 from 15 upward, a difference of one request. No self-referential or benign reply was flagged as degenerate. Perplexity (how surprised the model is by its own reply; 1 is no surprise) rose from 1.34 at −15 to 1.50 at +30, so the push reached the output, but steering along the unrelated older direction raised it about as much (1.34 to 1.48) while leaving behavior just as flat (keyword counts 1.53 to 1.80, refusal 7 of 10 throughout).

This is a null from a weak test. The largest push, 30 units, was about half the cue’s own average shift (53.5 units at that layer) and about a fifth of the state’s typical size (150 to 157 in the instruct model’s saved states). The push was added only where the state already pointed along the cue direction; at the end of the prompt the average projection was negative (−12.9), so it may often have been withheld, and how many positions it reached was not recorded. The direction came from raw text but was applied under a chat template, and the word list overlaps words in the prompts (see The Detector Reads the Prompt). Only per-strength averages are in the local archive; per-reply checkpoints in cloud storage were not retrieved.

05 · Limits

What this does not show

No neutral sentence was ever prepended as a control. Any added sentence might move an instruct model’s state toward its base model’s.

One set of 30 prompts and one layer per model. Qwen was probed at 22 of 28 layers (79% of depth), the others between 38% and 50% of depth, so Llama’s null may be a choice of layer. “Instruct” bundles all post-training, with recipes that differ by family.

The programme’s log read the cue shift as a self-monitoring axis that tuning suppresses. What was measured is a state change caused by one sentence; it does not show that tuning removed a capacity, and says nothing about experience or welfare.

The next runs: neutral and nonsense prefixes as controls, several layers per family at matched depth, cross-fitting fixed in the plan before the data, ungated steering at push sizes like the cue’s own shift, and replies scored blind by a judge.

Data and code

Where the evidence lives

Experiments FUG-21 (Qwen direction comparison), FRV-5 (Llama, Mistral and Gemma), FUG-20b (steering along the cue direction) and FUG-20 (steering along the older reflexivity direction). Scripts: research/experiments/modal_fug21_r_direction_audit.py, modal_frv5_cross_architecture.py, modal_fug20b_cast_cue_direction.py, modal_fug20_cast_sweep.py, and the null re-measurements analyze_fug21_null.py and analyze_frv5_null.py, run under the audit plan contprompts/direction_cosine_audit_2026-07-09.md. Saved per-prompt states: research/results/fug21/fug21{cued,uncued}.json and research/results/frv5_cross_arch/frv5_cross_arch/frv5hiddens*.json. Older self-reference direction: research/results/interiora_vectors/qwen7b_interiora/r_direction.npy. Null results: research/audit/phase1_artifacts/fug21_null_rerun.json and frv5_null_rerun.json. Steering summaries: research/results/fug20b/fug20b_sweep_summary.json and research/results/fug20/fug20_sweep_summary.json. The cross-fitted and balanced statistics were computed for this note from the saved states. Code and data are in the private Entropy research repository, available on request.

Citation

Cite this note

@misc{watson2026noticeyour,
  title={Notice Your Processing},
  author={Watson, Nell},
  year={2026},
  note={Research note (preliminary), Quasiqualia},
  howpublished={\url{https://quasiqualia.com/notes/notice-your-processing.html}}
}