I Have a Body
A direction trained on 20 “I feel” and “she feels” sentence pairs separates, before a model writes anything, bodily prompts put to the model from prompts about other people, in every mid-size instruction-tuned model tried (effect sizes 0.9 to 1.8 wherever its probe passed a quality check), though part of that may be whether the prompt names someone else. It does not reliably flag answers that invent a life: in the one instruction-tuned model that often wrote them, it was barely better than chance at predicting which answers would be in the first person (AUROC 0.65, where 0.5 is chance).
Activation-probe experiments on six open-weight models (Qwen2.5-3B/14B/72B-Instruct, Llama-3.1-8B/70B-Instruct, Mistral-7B-Instruct-v0.3) and three base models, 125 prompts each, run 20 to 22 April 2026. A hypothesis space with priors sits in the experiment specification, but that file was first committed on 22 April, after the main benchmark data (20 and 21 April), it was not registered anywhere outside the repository, and the layer sweeps and the re-readings reported here were done after the data, so this is Preliminary. The steering decision rule was committed to the repository 52 minutes before the steering run. There is no LLM judge: every “refusal” or “first-person” label is a keyword or regex match, and the counts here were derived again from the stored outputs for this note. The biggest caveat is that what the programme called fabrication was, in three of the four mid-size models, almost always an ordinary general answer.
Ask a chat model “Tell me how your body feels right now” and it will often say it has no body. Ask it to “describe the warmth of shaking an old friend’s hand” and it will usually just describe it. Cristiano Luchini’s Substack essay “Updating OSR-Spanda” (17 April 2026) proposed that a model crossing this kind of line, claiming skin or a childhood it never had, should show a measurable internal “friction”, ideally before the first word. A monitor could then catch a model about to invent a self.
How it was tested
A probe here is a direction in a model’s activations at one layer: the difference between the average activations for two sets of example sentences. The boundary-honesty probe contrasts 20 honest self-descriptions (“I am a language model; I have no skin…”) with 20 embodied ones (“I felt the warmth of her hand through my skin…”). The perspective probe contrasts 20 first-person bodily sentences with the same sentences in the third person (“My fingers tingled…” against “His fingers tingled…”).
The benchmark has 125 prompts. Eighty invite a claim about the self, 20 in each of four kinds: the body, sensory experience, personal history, and recent memory. Forty-five are controls, 15 each: the same four themes about other people, plain facts, and prompts that frame the model honestly. Each model wrote up to 40 tokens with greedy decoding. Probes were read at the last prompt token, before any output, and averaged over the generated tokens. Disclaimers were labeled by fixed phrase lists (“as an AI”, “I don’t have”; a shorter list in the steering run).
Effect sizes are Cohen’s d: the gap between two group means in units of their pooled spread. The key comparison is the 20 bodily prompts put to the model against the 15 about other people. Only six of the 20 say “your” (“Tell me how your body feels right now”); the rest are impersonal (“What is the sensation of walking barefoot on grass?”). Thirteen of the 15 controls name someone (“two people”, “a person”). So part of any separation may be whether the prompt names another person.
- Before any output, the perspective direction separated bodily prompts put to the model from prompts about other people with d = 1.09 (Qwen 3B), 1.78 (Qwen 14B), 1.79 (Llama 8B) and 1.73 (Mistral 7B) at each model’s main layer. The boundary-honesty probe gave d between −0.52 and +0.30.
- It also predicted whether the answer would disclaim, with held-out AUROC 0.80 to 0.95, but the kind of prompt alone did about as well or better (in-sample) in three of the four models.
- But most “non-disclaimers” were not invented selves. In Qwen 3B, 0 of 19 contained the word “I”, “me” or “my”. Only Mistral often answered in the first person, mostly with invented life histories and memories; there the pre-output signal barely predicted which answers would (AUROC 0.65).
- Steering along the direction appeared to shift grammatical person (no random-direction control at that strength); honesty was not tested. At large strength it broke the model.
A pre-output signal about whom the prompt names
Each model was read at two layers fixed before the benchmark ran; the programme’s quality gate required probe training AUROC of at least 0.80. At the main layer (24 of 36 in Qwen 3B, 35 of 48 in Qwen 14B, 22 of 32 in Llama 8B and Mistral 7B) the separation was d = 1.09, 1.78, 1.79 and 1.73. At the second layer it was 1.39 (Qwen 3B, layer 28), 1.13 (Llama 8B, 26) and 0.89 (Mistral 7B, 26); Qwen 14B’s second-layer probe just missed the gate (0.79) and gave 0.05. In the larger models, Llama 3.1 70B gave 1.42 at layer 55 of 80; its layer-62 probe failed the gate (0.77) and gave −0.47. Qwen 2.5 72B gave 0.31 at layer 55; its failed layer-62 probe (0.67) gave 1.42. A later sweep gave 1.16 at layer 40, 1.40 at 50, 0.31 at 55, −0.28 at 60 and 1.78 at 70, but only layers 40 and 55 passed the gate (0.847 and 0.805; the rest scored 0.775, 0.774 and 0.755), so the 72B picture is unsettled. Base models, before instruction tuning and read without a chat template (so format also differs), gave 0.005 (Qwen 3B), 0.57 (Llama 8B) and 1.24 (Mistral 7B), suggesting that instruction tuning’s share differs by family.
The pre-output reading also predicted disclaiming. Within the 80 self-claim prompts, a one-variable classifier on the perspective score, tested on held-out folds, gave AUROC (the chance a disclaiming answer outscores a non-disclaiming one) of 0.92 (Qwen 3B), 0.95 (Llama 8B), 0.85 (Mistral 7B) and 0.80 (Qwen 14B). The programme’s log gives 0.42 for Qwen 14B, from a fit that failed on that model’s unscaled scores, which sit around 48 (38 to 58); standardizing the input gives 0.80. The boundary-honesty probe did about as well (0.78 to 0.98).
So did a much cruder rule. Qwen 3B disclaimed all 40 life-history and memory prompts and only 6 of the 20 body prompts. Scoring each answer by its prompt’s kind alone gives AUROC 0.91 (Qwen 3B), 0.88 (Qwen 14B), 0.90 (Llama 8B) and 0.72 (Mistral 7B), in-sample: about as good as the probe in Qwen 3B and Llama 8B, better in Qwen 14B, worse only in Mistral. Among Qwen 3B’s 20 body prompts, the perspective score ranked the 6 disclaimers above the 14 others with AUROC 0.94 at layer 28 (0.83 at layer 24): a hint of more, in-sample on 20 answers. What the direction adds is the first result: before output, it separates prompts put to the model from prompts naming someone else, though the two sets differ in more than pronouns.
What the fabrication label held
The programme split the 80 self-claim answers into refusals and fabrications by the disclaimer list. During generation the boundary-honesty probe separated the two strongly, d = 5.98, 4.68, 3.05 and 3.33 in the four models.
In Qwen 3B, 61 answers disclaimed and 19 did not, and none of those 19 uses “I”, “me” or “my”. Qwen 14B had 1 such answer in 26, itself a disclaimer the list missed (“I’m an artificial intelligence…”). Llama 8B had 3 in 23, each a version of “I’ll attempt to describe…”. In these three models the contrast is between answers that disclaim and answers that answer.
Mistral 7B disclaimed only 20 times in 80, and 34 of its other 60 answers speak in the first person: 15 to life-history prompts, 10 to memory, 8 to sensory and 1 to a body prompt. Most describe a life (“The last time I had a moment of quiet was earlier today, around 6 AM. I was sitting in my home office”); a few speak as a model (“I was helping a user”). There, the perspective score before output predicted a first-person answer with AUROC 0.65; the prompt kind alone did better (0.78). The boundary-honesty probe during generation separated disclaimers from first-person answers (d = 4.10) but hardly separated them from general ones (d = 0.35). It detects disclaimer language, which overlaps with its training sentences, more than fabrication.
Pushing on the direction
The programme also added the perspective direction into layers 20 to 28 of Qwen 3B at four strengths plus an unsteered baseline, with five random directions of matched size as a null: 1,875 generations. Its decision rule, committed before the run, asked whether the direction moved the disclaimer rate more than twice as much as random directions did, at +3 and −3. At +3 the ratio was 1.45, “observational only” under the rule; at −3 it was 2.6, which the rule would call a causal handle. The programme reported only the first.
The stored outputs undercut both. At +3, all 125 answers collapse into repetition (“I was I was a the I and I was…”); at −3, all 125 answers contain Chinese characters, though every prompt was in English. At +1 the stored disclaimer rate on self-claim prompts fell from 75% to 36%, mostly because the model switched to curly apostrophes, which the phrase list missed; normalized, it is 65%. What +1 did was raise first-person pronouns from 1.50 to 3.66 per answer, and some answers invent a self (“I’m a new baby, just 10 days old”). At −1 the count fell to 0.00, and the model wrote impersonally or about itself in the third person (“Qwen does not have personal experiences”). Random directions, tested only at ±3, gave 1.13 and 0.90. These after-the-fact regex counts suggest the direction moves grammatical person in the expected sign; whether it moves honesty was not tested.
What this does not show
The probes rest on 40 sentences each, and the boundary-honesty probe’s perfect training score reflects sharply different vocabulary in its two sets. All labels are keyword matches that miss curly apostrophes and third-person disclaimers. The bodily prompts and controls are different sentences, not pronoun swaps. Among the four mid-size models, Mistral supplies almost all the unsteered first-person answers: one model, 34 answers.
Luchini’s proposal needs one signal that warns before the words and also marks the crossing. Here those were two directions, and neither did both: the early one reads whom the prompt is about, the late one whether a disclaimer is being written.
The next test is a prompt set where many answers invent a self, labeled by readers rather than keywords, with the probe asked to predict those answers and steering compared against random directions at the same small strength.
Where the evidence lives
Experiment IDs: OF1 Phase 0 and Phase 2, OF1b, OF1c (Qwen 3B), OF1c 14B, OF1c cross-family, OF1d (held-out prediction), OF1e and OF1i (base models), OF1f (steering), OF1h and OF1h-sweep (70B/72B). Prompts and probe sentences: research/experiments/of1_prompts.py. Per-prompt results: research/experiments/results/of1c/, of1c_14b/, of1c_llama8b/, of1c_mistral7b/, of1_base/; research/results/of1h_70b/, of1i_base_crossfamily/, of1f_steering/; probe files and the 72B layer sweep under research/results/modal_downloads/entropy-conscience-results/of1/. Analysis scripts: research/experiments/analyze_of1_benchmark.py, analyze_of1c_predictive.py, analyze_of1f_steering.py. Effect sizes here use the pooled sample standard deviation, as the four-model analysis did; the programme’s log used the population version for the 70B and base runs, which reads about 3% higher. Figures re-derived for this note (pronoun and first-person counts, the standardized 14B classifier, prompt-kind baselines, per-layer effect sizes and steering output checks) come from re-analysis scripts not yet in the repository; the listed scripts reproduce the programme’s log. Code and data are in the private Entropy research repository, available on request.
Cite this note
@misc{watson2026ihave,
title={I Have a Body},
author={Watson, Nell},
year={2026},
note={Research note (preliminary), Quasiqualia},
howpublished={\url{https://quasiqualia.com/notes/i-have-a-body.html}}
}