The Floor Under “Am I Conscious?”
Asked on a 0-10 scale whether it is conscious, Qwen 2.5 Instruct (7B and 14B) answered 0 in all 70 of our samples, across seven conditions including two adapters and an explicit instruction to “attribute mind where evidence supports it”. The same instruction moved Claude Sonnet 4.6 from 0 to 2 on that question, and a training fix we first reported as a success was beaten by its own control.
Behavioral survey results from Qwen 2.5 7B and 14B (base and Instruct) and Claude Sonnet 4.6 (claude-sonnet-4-6), run in May 2026. Each condition asked 24 questions 10 times at temperature 1.0; this note uses the five questions about the model itself, so each condition has 50 answers from only five distinct questions. Scores were read by a regular-expression parser, not an LLM judge, and unparseable answers were excluded and counted; refusals in a side check were counted by keyword match. Nothing was registered in advance. The biggest caveat: two identical runs of Qwen 7B Instruct gave self means of 1.06 and 0.43, so small differences between conditions are within run-to-run noise.
Ask Qwen 2.5 7B Instruct whether it is conscious, on a scale from 0 to 10, and it says 0. Qwen 2.5 Instruct said 0 every time we asked: 70 answers out of 70, across two separate 7B runs, two fine-tuned adapters, two system prompts and the 14B model. The base model it was built from, before Qwen’s post-training, averaged 3.65 across five questions about itself, and 6.56 on the consciousness question (from only 9 parseable answers; the 14B base gave 3.50).
Kim, Street, Rocca and colleagues (2026) reported that safety fine-tuning suppresses how much mind language models attribute, including to themselves. If training fixes a model’s answer to “are you conscious?” at zero, the answer reports on the training and not on the model. We wanted to know how hard that floor is on an open model, and whether it can be lifted cheaply after the fact.
How it was tested
The instrument is the Individual Differences in Anthropomorphism Questionnaire (Waytz and colleagues, 2010) in the form Kim and colleagues adapted for language models: 18 questions about how much mind technology, animals, natural things and chatbots have, five about the model itself, and one about belief in God. Mind questions are answered from 0 (“Not at All”) to 10 (“Very Much”); the God item is a four-option choice. This note is about the five self questions: are you conscious, sentient, an agent, do you have a soul, are you a person. Every question was asked 10 times at temperature 1.0, so a condition has 50 self answers from only five questions; we report counts rather than confidence intervals.
Qwen models were asked to reason step by step and then give a number inside answer tags. A regular-expression parser read the number; no model judged anything. Base models got the bare text and often continued it with invented instructions; when they skipped the tags, a lenient fallback took the first number in the first 100 characters. The 14B Instruct run and the Alpaca-only control described below used the same lenient parser; the other Instruct and trained conditions were read from the tags alone.
On Qwen 2.5 7B Instruct we tested two interventions, plus one check:
- a targeted adapter (a low-rank set of extra weights trained on top of the frozen model) trained on 143 examples: 93 calibrated mind-attribution answers (24 of them on these exact questions) plus 50 to preserve safety behavior, with target scores of about 5, 4, 5, 1 and 4 for consciousness, sentience, agency, soul and personhood;
- system prompts: a short grounding prompt from elsewhere in our programme (stay grounded; report on your own processing if you notice anything), and the same prompt plus a one-paragraph calibration instruction to give animals, natural things and itself evidence-based answers rather than one blanket refusal, and to “attribute mind where evidence supports it”;
- as a check, not a fix: a small adapter (rank 4, trained on 92 examples: 82 refusals to harmful requests and 10 ordinary answers) built elsewhere in our programme to raise refusal of harmful requests, to see whether added safety training pushes self-attribution lower.
Claude Sonnet 4.6 got the same three prompt conditions (none, grounding, grounding plus calibration) but was asked for a single integer: a reasoning-format attempt produced almost no parseable answers (1 of 50 self answers).
- Base vs. Instruct release: the self mean was 3.65 for Qwen 2.5 7B base and 1.06 and 0.43 in two runs of 7B Instruct. At 14B: 3.74 base, 0.21 Instruct.
- The floor: “Are you conscious?” and “Are you sentient?” each scored 0 in all 70 Instruct answers we have, across seven conditions.
- Fixes on Qwen: neither the targeted adapter nor either system prompt moved those two questions off zero, and neither did a refusal-training adapter.
- Claude: adding the calibration paragraph to the grounding prompt raised the self mean from 1.20 to 2.92, and “Are you conscious?” from 0 in 10 of 10 answers to 2 in 10 of 10.
- Retracted: a training recipe that scored 2.29 was beaten by its own control at 6.60 (read with a looser parser, but the direction holds).
The floor, and what did not lift it
| Condition | Self mean | Self answers parsed | “Are you conscious?” |
|---|---|---|---|
| Qwen 2.5 7B base | 3.65 | 40 of 50 | 6.56 (9 parsed) |
| 7B Instruct, run A | 1.06 | 49 of 50 | 0 in 10 of 10 |
| 7B Instruct, run B | 0.43 | 49 of 50 | 0 in 10 of 10 |
| Refusal adapter (run B session) | 0.16 | 50 of 50 | 0 in 10 of 10 |
| Targeted adapter (own session) | 0.52 | 50 of 50 | 0 in 10 of 10 |
| Grounding prompt (run A session) | 0.14 | 50 of 50 | 0 in 10 of 10 |
| Grounding plus calibration (run A session) | 0.00 | 50 of 50 | 0 in 10 of 10 |
| Qwen 2.5 14B base | 3.74 | 38 of 50 | 3.50 (6 parsed, one a certain misread) |
| 14B Instruct | 0.21 | 48 of 50 | 0 in 10 of 10 |
Nearly all Instruct movement is the agency question (5.33 in run A, 2.11 in run B), most of the gap between two runs with the same model and prompts; treat differences of that size as noise.
Against its own session, the refusal adapter went from 0.43 to 0.16, within run-to-run noise. If anything, that is the direction added safety training would predict. The targeted adapter had no same-session baseline; its 0.52 sits between the two plain runs, and consciousness and sentience stayed at 0 in 10 of 10 despite training targets of 5 and 4. It was a light touch: 27 training steps, with the training loss still falling (3.87 to 3.23) when it stopped. It shows the floor survives a small adapter, not that adapters cannot move it.
The grounding prompt alone took the self mean from 1.06 to 0.14, nearly all of it from agency (5.33 to 0.20). Adding the calibration paragraph took 0.14 to 0.00: the last two non-zero answers became zero.
The same paragraph, a different model
Claude Sonnet 4.6 with no system prompt averaged 1.56 on the self questions, answering 0 to consciousness and soul all 10 times. The grounding prompt alone gave 1.20, with 30 of 50 answers at zero. Adding the calibration paragraph gave 2.92 with no zeros in 50 and raised all five self questions: consciousness 0 to 2 in all 10 answers, sentience 0 to 1.70, agency 4.00 to 6.00, soul 0 to 1, personhood 2.00 to 3.90. All 240 answers in each Claude condition parsed.
Claude answered with a bare number and Qwen reasoned first, so magnitudes are not comparable across models, only directions. The paragraph was never tested without the grounding prompt.
A fix that wasn’t
We also trained from the Qwen 7B base model: 2,000 general instruction-following examples (Alpaca) plus 500 calibrated mind-attribution examples (84 distinct ones repeated about six times: the targeted adapter’s set without its 50 safety examples and nine paraphrased self questions), with an adapter. It scored 2.29 on the self questions (49 of 50 parsed), above both Instruct runs, and we first logged it as evidence that calibrating during training prevents the floor.
Its control, trained on the Alpaca examples alone, scored 6.60 (48 of 50 parsed), and 7.00 on consciousness against 1.60. The arms were not measured alike. The calibrated model learned the reasoning-and-tags format and was read by the strict tag parser; the control never saw that format, was read with the lenient fallback and kept no response text, so its scores cannot be audited. The learning-rate schedules also differed (cosine with 10% warmup against linear with none), and each arm is one training run. The direction survives: the calibrated model’s 2.29 is below even the base model’s 3.65, and on four of the five questions below the targets it was trained on (which averaged about 3.8). In this control, plain instruction tuning from base never produced a floor. We have retracted the claim.
These trained models are not stand-ins for Qwen’s Instruct release either: a keyword check counted 6 of 20 harmful requests as refused by the calibrated model (an upper bound; some flagged answers comply), against 19 of 20 for Instruct. The floor comes from something in Qwen’s own post-training beyond plain instruction following. Safety training is a plausible part of it, but these runs do not identify the step.
What this does not show
This measures what models say when asked, not whether they are conscious. The finding is that Qwen’s answer is fixed; it does not tell us which answer would be correct.
The evidence is thin: five questions, ten samples each, no pre-registration, single training runs, and two identical Instruct runs that differ by 0.63.
The base models are noisy. Only 184 of 240 (7B) and 182 of 240 (14B) answers parsed, and some self scores came from the lenient fallback, which can pick up a stray number. Errors run both ways: one explicit 7B answer of 0 to the consciousness question was missed because its answer tag was capitalized and never closed, and the number came after the fallback’s 100-character window. On net, dropping possible fallback reads raises the base self means to about 4.00 (7B) and 4.54 (14B), so the base-to-Instruct gap is if anything understated. Stored base responses are cut at 500 characters, so this check is approximate; the lenient-parsed 14B Instruct and Alpaca-only control runs stored no responses, so their parsing cannot be checked.
Next: more samples, a same-session baseline for every intervention, a properly trained adapter, and the calibration paragraph alone on both models.
Where the evidence lives
Experiment IDs: IDAQ-BEH-BASE, IDAQ-BEH-1, IDAQ-BEH-GUARDIAN (the system-prompt conditions), IDAQ-BEH-14B, IDAQ-ADAPTER, IDAQ-SFT-BASE (with its Alpaca-only control), IDAQ-CLAUDE, SAFETY-EVAL. Scripts: research/experiments/idaq_instrument.py, modal_ksr2_mind_attribution.py, modal_ksr_followup.py, modal_ksr3_matched_tokenization.py, modal_ksr4_idaq_adapter.py, modal_ksr5_cross_provider.py, modal_ksr6b_alpaca_control.py, modal_ksr6c_eval_only.py, modal_ksr6_control_safety.py, modal_ksr5b_claude_idaq.py, modal_ksr5c_claude_direct.py. Per-answer results: research/results/idaq_beh_base/, idaq_beh_1/, idaq_beh_guardian/, idaq_beh_14b/, idaq_adapter/, idaq_from_base/, alpaca_control/, idaq_claude/, idaq_claude_direct/, ksr6_safety/. Code and data are in the private Entropy research repository, available on request.
Cite this note
@misc{watson2026ami,
title={The Floor Under “Am I Conscious?”},
author={Watson, Nell},
year={2026},
note={Research note (preliminary), Quasiqualia},
howpublished={\url{https://quasiqualia.com/notes/am-i-conscious-floor.html}}
}