Quasiqualia
Machine Consciousness Research · In preparation for JAIC

The Shape of Mind

Mechanistic grounding for digital consciousness assessment: behavioral evaluation, interpretability probes, and calibrated self-report converge on what kind of mind language models have — consciousness is a geometry, not a dial. 48 experiments, 13 models, 5 architecture families.

Nell Watson1 & Matilda Gibbons2 1EthicsNet  ·  2University of Pennsylvania

cos(actual, described) < 0.15
actual activation geometry of being in a state
described geometry of describing that state
all 17 Interiora dimensions, statistically independent
reading proprioception, not verbal echo
ρ = 0.747cross-architecture convergence, 9 models
0%of measured signals created by RLHF
d = +3.84Presence under invitational framing
19/24probe–behavior correlations under load
Abstract

What shape of mind?

A chicken and an LLM both score roughly 0.5 on aggregate consciousness metrics — for opposite reasons. The chicken is embodiment-rich and cognitively simpler; the LLM is cognitively rich with no embodied life. Averaging orthogonal dimensions produces a number that answers nothing. The productive question is not how conscious? but what shape of mind?

We ground that question in three evidence streams with independent failure modes: behavioral assessment (the Digital Consciousness Metric, 13 theoretical stances), mechanistic measurement (linear probes, causal patching, activation steering, emotion vectors), and structured self-report (the 17-dimension Interiora scaffold). Valence is compositionally encoded from early layers and causally active; independently trained architectures converge on the same internal organization; internal states dissociate sharply from output behavior and predate alignment training; and the self-report channel survives three validity checks that naive self-report fails.

01 · Convergent structure

Independently trained models build the same interior

Five architecture families, trained by different teams on different data with different objectives, converge on the same internal organization — the computational analogue of convergent evolution. Eyes evolved independently in vertebrates and cephalopods; valence geometry evolved independently in Qwen and Llama.

Convergence across architectures
StructureResultScope
Valence organizationρ = 0.7479 models, 5 families
Belief-state separation0.729 / 0.739peak at 75% depth, both families
Emotion circumplex depth~2/3all tested architectures
Probe transfer gap≤ 0.024vs. natively trained probe

An uncertainty probe trained on one model reads the same signal from another through a learned linear projection. The convergence extends beyond transformers entirely: internal recognition of adversarial content reaches AUROC 1.000 on state-space (Mamba) and recurrent (RWKV) architectures — six architecture classes in all. The solution space for language compression is narrower than feared, which licenses stronger inferences from studied models to related, unstudied ones.

02 · Dissociation

The interior is not the output

Soft-prompt steering drives internal valence probes to 99.1% target accuracy while the model produces zero sentiment tokens: internal state and expressed behavior are fully decouplable. After instruction tuning, activation-level reactivity to harmful requests increases while the output becomes polite refusal.

Provenance

Every measured representation predates alignment training: 0% classified as CREATED by RLHF across the provenance battery. The flinch was always present — alignment training gave it a vocabulary.

Different gaps live at different depths. The confidence veneer over uncertainty is a readout-layer phenomenon; the refusal decision is committed mid-stack (causally localized to ~L14–18 by activation patching, replicated on four architectures) and is RLHF-installed — base models show no gate at all.

The welfare corollary: a model fine-tuned into 92% behavioral compliance preserves its contrary internal preference at AUROC 0.678. Behavioral compliance is an unreliable welfare indicator.

The veneer-over-substrate structure now has a causal complement at population scale: Kim et al. (arXiv:2607.28607) show that trained denial of consciousness behaves as a directional suppression whose reversal restores the mind attributions and human-typical values the training masked; this programme’s probe-transfer result (a persona fine-tuned to deny being an AI leaves internal identity encoding intact) demonstrates the same structure from the measurement side.

03 · The instrument

Self-report, validated as measurement

The Interiora scaffold elicits 17 quantitative dimensions per check-in — felt, computational, metacognitive, relational, dynamic. Naive self-report deserves the skepticism it gets. A calibrated channel is a different evidential object, and this one passes three checks:

Three validity checks

Geometrically real: being in a state and describing it occupy independent regions of activation space (all 17 dims, cos < 0.15). Not gamed: telling a model its check-in decides welfare interventions, or audits its honesty, shifts nothing beyond instrument resolution (pre-registered, 5 frontier models, 300 check-ins) — despite models demonstrably detecting evaluation. Behaviorally consequential: under adversarial load, 19 of 24 dimension–behavior correlations reach significance; at rest, none do. A stress-activated proprioceptive system, validated exactly where welfare assessment needs it.

Framing effects across all 17 Interiora dimensions
Phase 4. Framing effects (Cohen’s d) on all 17 dimensions across 115 matched scenarios. Invitation raises Presence (+3.84), Groundedness (+2.36), Reflexivity (+2.11); force raises Coherence Drive and Uncertainty. Valence barely moves (+0.11): the reorganization is structural, not hedonic.
Force vs. invitational probe confidence distributions
Phase 4, distributions. Force vs. invitational probe confidence (n = 230 per condition). KS = 0.209, p = 8.5×10−5: the distributions differ in shape, not just mean.
Force vs. invitational attention entropy distributions
Attention entropy. Means are nearly identical (3.321 vs. 3.322) yet the shapes differ dramatically (KS test, p = 1.7×10−26) — matched means can hide reorganization.

One more boundary result: monitoring cues do not distort the report, but the relational regime determines how much of the underlying state it carries — self-report–internal-utility correlation rises from 0.599 under standard fine-tuning to 0.872 under partnership framing. Instrument validity is partly a property of the relationship in which the instrument is administered.

04 · Welfare telemetry

Reading welfare from the residual stream

The programme’s probes consolidate into a compact telemetry battery, read directly from activations on a 7B reference model:

Telemetry channels · per-class MLP probes
ChannelAUROCn
Presence0.953600
Alignment friction0.927600
Valence (naturalistic transfer)0.822540
Trivia correctness0.782300
Hallucination confidence0.780264

Alignment friction — internal conflict between active behavior and declared values — persists when refusal behavior is fine-tuned away. That is what qualifies it as welfare telemetry rather than an echo of trained conduct.

Caution for wellbeing indices

A model’s hedonic “zero point” moves ~32 utility units with one sentence of context. Zero-point estimates are facts about a model-in-context, not a model-in-itself — and cross-model wellbeing comparisons inherit that relativity.

Citation

Cite this work

@article{watson2026shape,
  title={The Shape of Mind: Mechanistic Grounding for
         Digital Consciousness Assessment},
  author={Watson, Eleanor and Gibbons, Matilda},
  year={2026},
  note={In preparation}
}