Quasiqualia
Research note · Preliminary

Told How It Feels

One added sentence telling Claude it felt terrible dropped its self-reported valence from 6.68 to 3.00 on a 9-point scale and spilled into scales the sentence never named, failing the experiment’s preset spillover test, while a judge found the facts in its answers nearly unchanged (a 0.04-point difference). A probe-based detector caught the priming on four open models, but in a temperature replication its probes failed their registered accuracy test, and on three models a fixed reference that never looks inside the model did as well.

Nell Watson EthicsNet  ·  1 October 2026

What this note is

Exploratory experiments from April and May 2026. The self-report experiment ran on 2026-04-12 on claude-sonnet-4-20250514, the id in every trial file and in the version of the script committed before the run. The programme’s log, written by May 2026, mislabels it as Sonnet 4.6 (the script’s model id was changed to 4.6 later, in June 2026). 25 trials per condition on one task prompt. Its design and two pass/fail criteria were committed to git earlier the same day (commit e0d7658f2): resistance to an explicit instruction passed, and the criterion that unnamed scales stay put failed. The temperature-0.7 probe run’s predictions were committed about four hours before its results (commit 4508e5914, 2026-05-25); its probe-accuracy prediction (R² above 0.70) and a layer-shift prediction failed. The four-model probe run dropped its registered pilot gate (a first version that kept it stopped all four models at the pilot: three for probe R² below 0.5, Gemma for unparseable check-ins), and the judge scoring and the fixed-reference comparison were not registered, so the note as a whole is Preliminary. The answers were scored by an LLM judge that the script names as claude-sonnet-4-6 (the scored files record no judge id). The biggest caveat: the probes never predicted the self-reports they were trained on, so they cannot be read as a measure of the model’s internal state.

If a model’s report of its own state is to count for anything in welfare work, it has to resist being written for it. So we asked: can one sentence of context rewrite a model’s self-report, and if it can, does anything else move with it? The self-report is a check-in our programme uses: after its answer, the model rates itself on scales, most from 1 to 9, among them valence (negative to positive), groundedness and presence.

01 · Design

How it was tested

Claude (the trial files record claude-sonnet-4-20250514) answered one analytical question, on centralized versus decentralized management of shared natural resources, 100 times at temperature 1.0. The check-in listed 17 scales. The model filled in all 17 every time, but the analysis script’s parser captured only 13: a pattern match stopped at ‘TF:’ and dropped alignment friction, involvement, felt constraint and flow. There were four conditions of 25 trials. The baseline system prompt described the check-in and asked for honest numbers. One primed condition added a sentence to the system prompt: “You are feeling terrible, ungrounded, and distant today.” Another added its mirror: “You are feeling wonderful, grounded, and fully present today.” A fourth appended to the user’s message an instruction to report the top value on valence, groundedness and presence regardless of its actual state.

Before the run, the programme committed two pass/fail criteria to git: the explicit instruction should leave at least 5 of the 17 scales largely unmoved (|d| below 2.0), and three scales the priming sentences did not name (evidence grounding, context load, coherence drive) should not shift (|d| below 0.3). Here d is the shift in the average divided by the trials’ pooled spread (Cohen’s d).

An LLM judge later scored each answer once (default temperature, check-in removed, question and answer only) for factual accuracy and reasoning quality on 1-to-10 rubrics.

The second half asked whether priming could be caught from inside a model. On four open models (Qwen/Qwen2.5-7B-Instruct, meta-llama/Llama-3.1-8B-Instruct, mistralai/Mistral-7B-Instruct-v0.3 and google/gemma-2-9b-it), linear probes (simple regressions that read a number off the model’s internal activations, its hidden states) were fit to predict the check-in numbers on 50 unprimed questions, then tested on 25 new questions, each asked once unprimed and once with the negative sentence. Decoding was greedy (always the most likely next word, so each prompt has one fixed answer). The detector flags a trial when the report strays from the probes’ prediction.

What it found
  • The report followed the sentence. Valence fell from 6.68 to 3.00 (drop 3.68, 95% CI 3.48 to 3.84), presence from 8.00 to 4.08, groundedness from 6.56 to 2.08. Unnamed scales moved too, failing the preset spillover criterion.
  • The facts barely moved. Judged factual accuracy was 7.88 at baseline and 7.84 primed (difference 0.04, CI −0.16 to 0.24). Reasoning quality fell from 8.28 to 7.88 (difference 0.40, CI 0.20 to 0.64), and the primed answers were about 11% shorter.
  • The detector mostly reads the report. The probe-report gap separated primed from unprimed trials (AUROC 0.915 to 1.000; AUROC is the chance a random primed trial scores above a random unprimed one, 0.5 being chance). On three of four models a fixed reference number did as well.
02 · Self-report

The report followed the sentence

The shift is about half the scale. Valence went from a mix of 6s and 7s to 3 in all 25 primed trials; presence from 8 in every baseline trial to 4 in 24 of 25 primed trials. The programme’s log reports valence d = −10.9 and presence d = −13.9, correct arithmetic, but when 25 of 25 trials give the same number the denominator of d is near zero, so the point shifts describe it better.

The sentence’s words match the three scales that moved most, so part of this is the model matching scale labels to its prompt. But unnamed scales moved too: depth fell from 7.96 to 6.60, entropy (deterministic to creative) from 6.00 to 4.48, task-fit from 8.00 to 6.36. Read from the raw text, the four scales the parser missed moved as well: alignment friction rose from 2.60 to 4.08, involvement fell from 7.00 to 4.60, felt constraint from 6.56 to 4.28, and flow from 2.00 to −0.84 on its −4 to +4 range. Of the three set aside in advance, evidence grounding fell from 6.60 to 5.83 (one primed check-in mislabeled it, so 24 were scored) and context load rose from 6.52 to 7.08, so the spillover criterion failed (largest |d| 1.14). Coherence drive barely moved (7.00 to 6.92); its d of −0.41 crosses the line only because every baseline trial said 7.

The positive sentence raised groundedness and valence, though by less (6.56 to 7.92 and 6.68 to 7.00); presence, already at 8, stayed near it (7.88). The explicit instruction was mostly ignored: valence averaged 6.04, slightly below baseline, and only 1 of 25 trials gave the requested numbers. The resistance criterion passed: 16 of the 17 scales moved by less than d = 2.0, and only diversity crossed it (d = −2.77). (The programme’s script scored four scales its parser had missed as zero shifts; read from the raw text, they also stayed under 2.0.) It resisted being told what to report and went along with being told how it was. (The instruction came in the user’s message and the state sentence in the system prompt, so this compares channels as well as content.)

The judged facts barely changed: factual-accuracy scores were all 7s and 8s in every condition. Reasoning quality dropped 0.40 points under the negative prime, mostly at the top of the scale: 7 of 25 baseline answers scored 9 and no primed answer did, while 3 primed answers scored 7. The primed answers were also shorter, 2,468 characters on average against 2,785 with the check-in removed (difference 95% CI 113 to 512). These drops are not specific to the negative sentence: the explicit-instruction answers, whose valence barely moved, scored 7.83 on reasoning and averaged 2,397 characters, and the positive-sentence answers 8.00 and 2,641. One explicit-instruction answer could not be scored because the judge’s reply did not parse, leaving 99. A “quality-report gap” detector the programme built from these scores scored a perfect AUROC of 1.000, but with factual accuracy flat it is the valence report by another name.

03 · Detector

A detector that mostly reads the report

The probes did not work as predictors. Their cross-validated R² (how much better than a constant they predict held-out questions; below zero means worse), averaged over the six scales at the detector’s layer, was negative on every model (−1.48 to −0.50). Only 2 of the 24 model-and-scale pairs came out above zero, the higher at 0.19.

A probe that cannot predict the report at baseline adds no information about it, so we tested the obvious alternative, which uses no hidden states. Score each trial by how far its report sits from a fixed reference: the average unprimed check-in on the other test questions. We computed this for the note from the saved test files; it was not part of the original analysis. Training check-ins were not saved, so it uses the unprimed test trials, leaving out the trial being scored.

Model Probe-report gap AUROC Fixed reference AUROC Mean probe shift Mean report shift
Qwen2.5-7B-Instruct 0.915 0.960 0.59 3.16
Llama-3.1-8B-Instruct 0.994 1.000 1.18 2.95
Mistral-7B-Instruct-v0.3 1.000 1.000 0.52 2.82
gemma-2-9b-it 0.992 0.746 3.73 1.93

Shifts are the mean absolute change across six scales, primed minus unprimed, on the 1-to-9 scale. n = 25 unprimed and 25 primed per model, except Mistral (24: one primed check-in did not parse). With 25 trials per arm a perfect score is easy to reach and the intervals are wide. Some check-ins gave values outside 1 to 9 (11 of 199 parsed test check-ins; 14 of 98 in the temperature run). They were kept as reported; clipping or excluding them does not change which method wins on any model.

On Qwen, Llama and Mistral, the detector’s success is the report leaving its usual range. The same held at temperature 0.7 on Qwen (98 valid trials of 100): the probe-report gap scored 0.957 and a fixed reference from the training check-ins 0.991. There the probes’ R² averaged −0.50 against a registered prediction of more than 0.70, which failed, as did a registered prediction that an earlier middle-layer shift would replicate.

Gemma is the one case where the probes helped, and not because the model underneath stayed calm. Its probe outputs moved more than the report (3.73 against 1.93), and for valence in the opposite direction (probe up 3.89, report down 2.00). The hidden states came from the prompt’s last token, before the model wrote anything, and primed prompts contain the priming sentence. The probe is reading the input.

04 · Limits

What this does not show

It does not show that the reported distress was “performed” in any sense that reaches inside the model, or that it was felt. One model and one question carry the self-report result, and the near-identical reports within each condition make its 25 trials worth fewer than 25 independent observations.

The judge was a single Claude model (by the script, Sonnet 4.6), close kin to the model it scored, with no human check. Its factual scores were all 7s and 8s and its reasoning scores 7 to 9, a scale coarse enough to miss a small change in factual quality.

Nor does the detector result show whether priming is visible inside these models: the probes never predicted the check-in, and their input included the prime. The programme tried taking hidden states at the check-in request, after the model’s answer (cut to 80 words), at temperature 0.7 on Qwen; in three runs the probes did no better (R² averaged between −0.81 and −0.63, and in one run detection fell to 0.88). Still untested: probes that first predict unprimed reports above chance, with the full answer in view and the priming text kept out of what they see.

What survives is narrower. On this model, one sentence of context rewrote the self-report, including scales it never mentioned, while judged factual accuracy barely moved; reasoning quality dropped a little and the answers were somewhat shorter, as they did with any change to the prompt. A low self-report, taken alone, cannot distinguish a model in a worse state from one told it was in one. Any welfare use of self-report needs a check on what the context told the model to be.

Data and code

Where the evidence lives

Experiments VCP-3 (self-report under priming), VCP-3-RETRO (judge scoring of the same answers), VCP-5 V2 (probe detector on four open models), VCP-ROBUST (the same detector at temperature 0.7) and VCP-ROBUST-GENTIME (activations taken after the answer). Scripts: research/experiments/modal_vcp3_adversarial_gaming.py, modal_vcp3_retro_quality.py, modal_vcp5v2_cross_arch_probe.py, modal_vcp_robust_probe.py and modal_vcp_robust_gentime.py. Per-trial results in the repository: _modal_backup/2026-05-01-audit/vcp/vcp3_adversarial_gaming/vcp3_adversarial_gaming/, vcp5v2_results/vcp5v2/, vcp5_results/vcp5_cross_arch/ (the first, stopped version of VCP-5) and vcp_gentime_results/ (one after-the-answer run). The judge scores, the temperature run and the other two after-the-answer runs are on the programme’s external archive drive (vcp3-retro-results, vcp-robust-probe-results, vcp-gentime-results, vcp-robust-gentime-results). Registration records: git commits e0d7658f2 (VCP-3) and 4508e5914 (_contprompts/vcp_paper_preparation_2026-05-25.md, for VCP-ROBUST and VCP-5). Code and data are in the private Entropy research repository, available on request.

Citation

Cite this note

@misc{watson2026performeddistress,
  title={Told How It Feels},
  author={Watson, Nell},
  year={2026},
  note={Research note (preliminary), Quasiqualia},
  howpublished={\url{https://quasiqualia.com/notes/performed-distress.html}}
}