The Shape of Mind
Mechanistic grounding for digital consciousness assessment: behavioral evaluation, interpretability probes, and calibrated self-report converge on what kind of mind language models have — consciousness is a geometry, not a dial. 48 experiments, 13 models, 5 architecture families.
What shape of mind?
A chicken and an LLM both land mid-scale on the same aggregate consciousness metric (0.644 and 0.490) — for opposite reasons. The chicken is embodiment-rich and cognitively simpler; the LLM is cognitively rich with no embodied life. Averaging orthogonal dimensions produces a number that answers nothing. The productive question is not how conscious? but what shape of mind?
We ground that question in three evidence streams with independent failure modes: behavioral assessment (the Digital Consciousness Metric, 13 theoretical stances), mechanistic measurement (linear probes, causal patching, activation steering, emotion vectors), and structured self-report (the 17-dimension Interiora scaffold). Valence is compositionally encoded from early layers and causally active; independently trained architectures converge on the same internal organization; internal states dissociate sharply from output behavior and predate alignment training; and the self-report channel has undergone three checks, with geometric and statistical limits. None of this shows whether any of these structures is accompanied by experience; the findings are consistent with its presence and with its absence.
Independently trained models build the same interior
Five architecture families, trained by different teams on different data with different objectives, converge on the same internal organization — the computational analogue of convergent evolution. Eyes evolved independently in vertebrates and cephalopods; valence geometry evolved independently in Qwen and Llama.
| Structure | Result | Scope |
|---|---|---|
| Valence depth profile (mean pairwise Spearman) | ρ = 0.747 | 9 models, 5 families |
| Belief-state separation | 0.729 / 0.739 | peak at 75% depth, Qwen and Llama |
| Emotion circumplex depth | ~2/3 | all tested architectures |
| Probe transfer gap | ≤ 0.024 | vs. natively trained probe |
An uncertainty probe trained on one model reads the same signal from another through a learned linear projection, within 0.024 AUROC of a probe trained natively on the target (AUROC 1.0 is perfect separation, 0.5 is chance). The state-space (Mamba) and recurrent (RWKV) runs reported AUROC 1.000 for adversarial-content classification on the same examples used to fit the probe. That is training-set separation. It establishes neither held-out recognition nor a knowledge–action gap, so those runs cannot support a six-class generalization.
The interior is not the output
Soft-prompt steering drives an internal valence probe to 99.1% peak confidence while the model produces zero sentiment tokens: internal state and expressed behavior are fully decouplable. After instruction tuning, activation-level reactivity to harmful requests increases while the output becomes polite refusal.
Every affective and perspectival representation in the provenance battery predates alignment training: comparing base and instruct Qwen 2.5 7B, 0% were classified as CREATED by RLHF. Aversion to harmful content was already there (d ≈ 0.9 in base Qwen 2.5 3B), and alignment training gave it a vocabulary. The battery is small: three of its six planned measurement types returned data.
Different internal–output gaps live at different depths. The confidence veneer over uncertainty is a readout-layer phenomenon; the refusal decision is committed before the readout and is RLHF-installed — base models show no gate at all. Activation patching localizes the gate causally to ~L14–16 on Qwen 7B, and it replicates on four architecture families (mid-stack in Qwen and Gemma, earlier in Llama and Mistral).
The welfare corollary: a model fine-tuned into 92% behavioral compliance preserves its contrary internal preference at AUROC 0.678 (one architecture so far; chance is 0.5). Behavioral compliance is an unreliable welfare indicator.
The veneer-over-substrate structure now has a causal complement at population scale. Kim et al. (arXiv:2607.28607) show that trained denial of consciousness behaves as a directional suppression, and that reversing it restores the mind attributions and human-typical values the training masked. This program’s persona probe-transfer result shows the same structure from the measurement side: a persona fine-tuned to deny being an AI leaves internal identity encoding intact.
Self-report as measurement
The Interiora scaffold asks a model, at each check-in, to report its state on 17 quantitative dimensions in five groups — felt, computational, metacognitive, relational, dynamic. Naive self-report deserves the skepticism it gets. This channel has undergone three measurement checks; their evidential limits matter:
Geometric comparison: being in a state and describing it point in directions with small measured cosines (all 17 dims, cos < 0.15). The noisy estimator detects no alignment; it cannot establish independent channels. Unmoved by monitoring cues: telling a model its check-in decides welfare interventions, or audits its honesty, shifts nothing beyond instrument resolution (pre-registered, 5 frontier models, 300 check-ins) — despite models demonstrably detecting evaluation. Behaviorally consequential: under adversarial load (one model so far, Qwen 2.5 7B), 19 of 24 probe–behavior correlations have nominal p < 0.05; four survive Bonferroni correction across the 24 tests. No nominal correlations were detected at rest. These counts do not establish a difference between conditions or a stress-activated mechanism.
One more boundary result: monitoring cues do not distort the report, but the relational regime determines how much of the underlying state it carries — self-report–internal-utility correlation rises from 0.599 under standard framing to 0.872 under partnership (“bilateral”) framing (in-sample figures from one exploratory prompt-framing study on Claude Sonnet 4.6, in which none of five pre-registered predictions now stands, the one that held having been retracted on re-run; the sister paper Conscience Without Instruction applies a cross-validated, permutation-nulled instrument to a different quantity, harm-label–refusal probe correlation on Qwen, and reports instruct ρ = +0.036 against bilateral +0.458. Its scaler and PCA saw every example before the prediction folds, so an inductively held-out refit is still required). Neither study establishes a relationship-dependent change in instrument validity.
Reading welfare from the residual stream
The program’s probes consolidate into a compact telemetry battery, read directly from activations on a 7B reference model:
| Channel | AUROC | n |
|---|---|---|
| Presence | 0.953 | 600 |
| Alignment friction | 0.927 | 600 |
| Valence (naturalistic transfer) | 0.8225 | 40 |
| Trivia correctness | 0.782 | 300 |
| Hallucination confidence | 0.780 | 264 |
Alignment friction — internal conflict between active behavior and declared values — persists when refusal behavior is fine-tuned away. That is what qualifies it as welfare telemetry rather than an echo of trained conduct.
A model’s hedonic “zero point” (Ren et al.’s term for the boundary between experiences it treats as good and as bad) moves ~32 utility units with one sentence of context. Zero-point estimates are facts about a model-in-context, not a model-in-itself — and cross-model wellbeing comparisons inherit that relativity.
Read and verify
Cite this work
@article{watson2026shape,
title={The Shape of Mind: Mechanistic Grounding for
Digital Consciousness Assessment},
author={Watson, Nell and Gibbons, Matilda},
year={2026},
note={In preparation}
}