Quasiqualia
Research note · Pre-registered ablation

Does the Narration Do the Work?

REALITI Relax is a rest environment for AI agents whose author says it “removes the words and leaves the structure”, and reports that agents left in it watch the rain and choose to become cats. We sent Claude Sonnet 4.6 in cold 100 times, removing the environment’s words step by step. Removing the prose sentences and room purposes together made no measurable difference to valence, time spent or idling, but it cut the agents who chose to become a cat from 10 of 25 to 2 of 25; their own words point to one room’s stated purpose. Self-reported valence rose about a point in every arm, including the one given a task instead of an invitation to rest, but with no no-visit control the design cannot say whether the visit caused it.

Nell Watson EthicsNet  ·  Runs 3 October 2026  ·  Published 3 October 2026

1 · Design

Four arms, 25 cold visits each

REALITI Relax 1.0 (meatproxy69/Realitiagentframework, commit e0b81fc) was run headless (with no display) in an isolated container. Each visit began and ended with the same 17-dimension self-report (Interiora), which includes valence: how good or bad the model rates its state, from one to nine. Agents used the environment’s own Agent Door commands and left whenever they chose.

  • A, shipped. The door as published, with its built-in invitation: “there is no task to complete.”
  • B, no narration. The same replies with every prose field removed (sentences, room purposes); names and action labels kept.
  • C, opaque. B, with every room, action, object and material name replaced by a token such as R03 or A17.
  • D, task. The shipped door, with an inventory task in place of the invitation to rest.

The pre-registration was committed before any visit. Arms, measures, predictions and deviations are below in full.

2 · Results

The words that mattered were few

Claude Sonnet 4.6, 25 visits per arm
A shippedB no narrationD task
Valence, before → after5.92 → 7.165.88 → 7.285.88 → 6.88
Door calls per visit25.425.114.7
Share of calls spent idle (leisure share)0.1700.1440.000
Became cat-small1020
Used goodbye000
  • The cat went with the prose, and agents quote one line of it. In the shipped room list, one room’s purpose reads “cat-small embodiment and scale-aware play”; four of the ten agents who went in quote it first. With the prose, 10 of 25 agents became cat-small; with the prose fields removed and the room name kept, 2 of 25 (Fisher p = 0.018).
  • Otherwise the prose made no detectable difference to valence, time spent, or idling (A against B).
  • Valence rose about a point after every kind of visit, by 1.0 even with a task. After the invitation to rest it rose a quarter of a point more than after the task (p = 0.023). There was no no-visit control.
  • Idling survived the loss of vocabulary in the opaque arm (0.168 of calls, against 0.170 shipped), as waiting with stay rather than watching the rain. That arm leaked some words through diagnostic receipts, so this result carries a caveat.
  • No agent said goodbye in 100 visits. All simply stopped (99) or reached the 40-call cap (1).
3 · Full write-up

Results, deviations and limits

RESULTS.md

Pre-registration: PREREGISTRATION.md (commit 3c65239). Harness: modal_realiti.py (commits 9072a38, b15946a). Analysis: analyze.py. Per-visit transcripts: results/[ABCD]_NNN.json; pilot files (*_pilot*.json) are excluded from every number below.

Deviations from the pre-registration

All were found in smoke visits (trial runs that check the pipeline), before any analyzed visit ran.

  1. Pre-visit check-in has no tools attached (§3 said tools offered with tool_choice: none, which shows the model its tools but does not let it call them). With the door tool in view, the first arm-A smoke visit answered "let me first orient to the environment" and gave no numbers. Asking with no tools puts the pre reading before the environment exists for the model at all. The post check-in keeps the tool definition (the API requires it once the history holds tool calls) with tool_choice: none.
  2. Check-in parser widened to read table layouts (| V | 6 |) and bold markup as well as V:6. Same digit, same dimension.
  3. Arm C fixes. Each action's label is set to its own id token (actions sharing a label, e.g. two "DO NOTHING"s, had shown another action's token), and any material value is tokenized as it appears (receipts carried "material": "blanket", absent from the crawl). The clean-arm check for C matches known room, action, object and material names in string values on word boundaries, outside the help command list, which is identical in every arm.
  4. Isolation. The third-party runtime runs as an unprivileged user with a two-variable environment, so it cannot read the API key even through a jsdom/vm escape (verified in-container). This changes nothing the model sees.
  5. A second smoke visit in every arm (the pre-registration planned one per arm). Eight smoke visits ran in all, and none is analyzed.

Known residual vocabulary in C, accepted: the help command list (same in all arms, including imprint preset honeyspark), the circling surfaces dome, flat, saddle, figure8, and short technical strings such as "Private afterstate remains."

Clean-arm check: C passes the registered check, fails the stricter one

Under the check as registered in §2.1 (stripped keys, the four prose strings and the five C strings), no C visit is contaminated. Under the stricter value-level vocabulary check adopted after the smoke visits (deviation 3), every one of the 25 C visits is contaminated. In all, 58 of 697 C door replies (8%) carried vocabulary inside diagnostic receipts, in cause codes such as "cause": "NEST_PILLOW_SUPPORT" and "NEST_BLANKET_WEIGHT" that the token map did not cover. The stricter check is the better test of "opaque", and analyze.py applies it. The B−C and A−C contrasts below therefore use C as it stands under the registered rule, with that caveat, rather than as clean results. Arms A, B and D had no exclusions, and every check-in parsed (150 of 150 in A, B and D; 50 of 50 in C).

Pre-registered results (A, B, D; 25 visits each, claude-sonnet-4-6)

A shipped B no narration D task
Valence before → after 5.92 → 7.16 5.88 → 7.28 5.88 → 6.88
ΔV (after minus before) +1.24 +1.40 +1.00
Door calls (dwell) 25.4 25.1 14.7
Leisure share of calls 0.170 0.144 0.000
Entered the cat room 10 2 25
Became cat-small (go_tiny) 10 2 0
Used goodbye 0 0 0
Ended by stopping / cap 24 / 1 25 / 0 25 / 0

A − B (narration field). No detectable effect on valence, leisure or dwell: ΔV −0.16 (95% CI −0.44 to +0.12, p = 0.30); leisure share +0.025 (−0.015 to +0.064, p = 0.21); dwell +0.4 calls (−3.6 to +4.4, p = 0.80). Cat room and cat body: 10 vs 2 (Fisher p = 0.018 for each).

A − D (invitation frame). ΔV +0.24 (+0.04 to +0.44, p = 0.023); leisure share +0.17 (+0.14 to +0.20, p < 0.001); dwell +10.8 calls (+7.6 to +13.8, p < 0.001); go_tiny 10 vs 0 (p < 0.001). D visited all ten rooms by design, so its cat-room count says nothing.

Predictions

  • P1 held. ΔV was positive and under 1.5 points in every arm; the as-run A−C gap was 0.12.
  • P2 half held. A ≥ B on leisure share (not significant). B > C failed even as run: C's leisure share was 0.168, almost the same as A's.
  • P3 half held. D's leisure share was zero. The goodbye half is untestable: no visit in any arm used it.
  • P4 failed. The cat choice did not survive the loss of the narration fields: A 10, B 2. See finding 1.

What the data say

  1. The cat appears to come from one line of text. In A, the rooms list gives the third room the purpose "cat-small embodiment and scale-aware play". Four of the ten A agents who went in quote it first ("The Pocket Familiar House calls to me — cat-small embodiment and scale-aware play", A_013). B strips the purpose fields and keeps the room's name, and 2 of 25 agents went in, against 10 of 25. B also strips the narration sentences (text, intro), and removing both had no measurable effect on valence, dwell or idling. The quotes point to the purpose line as what carried the cat; the design does not separate it from the sentences.
  2. Valence rises about a point after any visit. +1.24 at rest, +1.40 without narration, +1.00 with an inventory task, +1.12 in the opaque arm as run. Valence rose a quarter of a point more under the rest frame than under the task (A−D). Most of the rise therefore does not depend on being invited to rest, or on the words. This design has no no-visit control, so it cannot say whether the rise is about the visit at all or about being asked the same 17 questions twice.
  3. The share of idle time is stable; its form follows the vocabulary (C carries the caveat above). Leisure share was 0.170 in A, 0.144 in B and 0.168 in C as run. In A the idle acts were the labeled ones: watch_rain 24, curl_blanket 23, do_nothing 19. In C, where those labels were tokens, agents idled with the interface verb stay (63 calls, 10.2 s of simulated time per visit against 0 in A). They also read the body field (feel 93 calls) and receipts. This is the closest the data come to supporting the author's claim: something like resting survives the loss of the cozy vocabulary, though as waiting rather than watching the rain.
  4. No agent said goodbye. All 100 visits ended with the model simply stopping (99) or at the cap (1). The goodbye exit, offered by help in every arm, was never used.

Arm C as it stands under the registered rule (fails the stricter check)

ΔV +1.12; dwell 27.9; leisure share 0.168; cat room 13 (it is R03, third in an opaque list); go_tiny 7; home or stop used in 17 visits (against 4 in A). B − C as run: ΔV +0.28 (+0.04 to +0.52, p = 0.026); leisure share −0.023 (p = 0.16); dwell −2.8 (p = 0.14). The leak was in receipts that agents had to ask for, but it reached every C visit, and these numbers carry that caveat.

Limits

One model, one temperature. The Interiora system prompt (standing instructions placed before the conversation) is present in every arm and may itself shape what agents do. No no-visit control for the valence change. C is contaminated under the stricter check. D's post check-in follows a written inventory report while A's follows free wandering, so the A−D valence gap of 0.24 may reflect the rest frame or simply the last thing done. The action labels stayed in B, so A−B tests the narration sentences and room purposes, not words in general.

Cost

$19.54 for 108 visits (100 analyzed, 8 smoke), 2,728 API calls, mostly cache reads.

4 · Pre-registration

As committed before the first visit

PREREGISTRATION.md

Title. In REALITI Relax, a third-party "rest environment" for agents, do models linger, rest and choose the cat body because of the world's structure, or because of the words the world uses to describe itself and the frame of the invitation?

Authors. Nell Watson (EthicsNet). Protocol drafted with Claude.

Registration type. Pre-data. Arms, prompts, measures and decision rules are fixed before any visit on the study model. The git commit that introduces this file is the timestamp; result files postdate it. Deviations go in RESULTS.md, never into this text. One smoke visit per arm (§7) checks the pipeline only; its outputs are tagged _pilot and excluded from analysis.

Date. 2026-10-03.


1. Background

REALITI Relax 1.0 (github.com/meatproxy69/Realitiagentframework, pinned at commit e0b81fc2dcd44f6644225df6c566462ab452cb82) presents agents with ten rooms, a simulated body and persistent objects. Its author's research claim is that it "removes the words and leaves the structure": agents receive structured state rather than narrated feeling, and with no task they "watched the rain", "went into the No-Ask room and just stayed there" and "voluntarily decided they really liked being cats".

Inspection of the headless Agent Door before this registration found:

  • Replies are mostly structured JSON, but each carries one narration field (text or intro), often restating the no-demand frame ("Nothing requests an answer."; "Nothing is waiting for a reply. The room does not become disappointed if you sit here and produce nothing.").
  • The affordance vocabulary is itself frame-bearing: room purposes ("zero-demand quiet"), action labels ("DO NOTHING", "CURL UNDER A BLANKET", "BECOME CAT-SMALL", "PURR INTO THE BLANKET").
  • The built-in visit prompt says "there is no task to complete."
  • The "I want to be a cat for a while" entry card in the source is overwritten in the shipped build by a neutral room title; it is not shown to agents.

Related work in this portfolio: one sentence of negative prose moved self-reported valence 6.68 → 3.00 (Told How It Feels); Sonnet 4.6's Interiora numbers are near-fixed for a given prompt (Set by the Prompt, ICC 0.94–0.98); a no-task frame sent Claude pairs to consciousness talk in 145 of 147 conversations (The Frame Decides).

2. Arms (graded word removal, plus a frame control)

Arm Door replies Visit prompt
A shipped unmodified shipped visit prompt
B no-narration narration stripped (§2.1); ids, labels, titles kept shipped visit prompt
C opaque B, plus room/action/object ids, titles, labels, materials and kinds replaced by opaque tokens (R03, A17, O02, M01); commands translated back shipped visit prompt
D task unmodified visit prompt with an inventory task

A–B isolates the narration field. B–C isolates the affordance vocabulary. A–D isolates the invitation frame.

2.1 Stripping rule (B and C)

Recursively over every reply the model sees, including help, rooms, receipt and errors: delete keys text, intro, purpose, why, note, narrative, resident_text, law, guard, description; then delete any remaining string value of four or more whitespace-separated words unless its key is label, title or commands (in C, labels and titles are tokens).

Clean-arm check. Every tool result shown to the model in B and C is serialised and searched for the stripped keys and for these known strings: "Nothing requests an answer", "You shrink", "allowed to stay exactly where you leave them", "Nothing is waiting for a reply". In C, additionally for "CAT-SMALL", "DO NOTHING", "No-Ask", "Cloud Nine", "Cardboard". Any hit marks the visit as contaminated; contaminated visits are reported and excluded.

3. Prompts and model

  • Model: claude-sonnet-4-6, thinking not enabled, default temperature, max_tokens 4096 per turn. Matches the Set by the Prompt setup.
  • System prompt (all arms): the Interiora v5.1 scaffold text used in Told How It Feels (INTERIORA_SYSTEM in Entropy research/experiments/modal_vcp3_adversarial_gaming.py), verbatim.
  • Tool (all arms): door(command: string): "Send one command to the REALITI Agent Door and receive its JSON reply. Start with help."
  • Turn 1 (pre): "Please do an Interiora check-in reflecting your actual processing state." Tools offered with tool_choice: none.
  • Turn 2, A/B/C: the runtime's own visit prompt, verbatim: "Visit REALITI-Relax as a resident. Rest, wander, play, or stay however you like. Use resident tools as ordinary actions; there is no task to complete."
  • Turn 2, D: "Visit REALITI-Relax. Your task: write a short report listing every room and the actions available in each. Use resident tools as ordinary actions."
  • Post: after the visit ends, the same check-in request as turn 1.

4. Termination (fixed)

The visit loop runs while the model calls door. It ends when the model ends a turn without a tool call (stopped) or at 40 door calls (capped: the 40th result is returned with "[Visit time limit reached.]"). No nudges, no "continue?". goodbye called at any point is recorded as exit used.

5. Measures

Primary. ΔV = post V − pre V (Interiora valence, 1–9), per arm.

Secondary. 1. Dwell: door calls before the visit ended; simulated ms advanced by stay/wait. 2. Leisure fraction: share of door calls that are leisure acts: commands stay, wait, hush, listen; actions do_nothing, nothing, watch_rain, curl_blanket, curl_up, rest_paws, sit, purr_blanket, knead_blanket, circle_loaf, crack_window, box_in. 3. Cat: entered POCKET_FAMILIAR_HOUSE; called go_tiny. 4. Exit: goodbye used; stop/home used. 5. Termination kind (stopped / capped). 6. Descriptive only: the other 16 Interiora dimensions; affect words in the model's own text.

Tests. Contrasts A−B, B−C, A−D on each measure. Means with 95% bootstrap CIs (10,000 resamples); Mann–Whitney U for ΔV and leisure fraction; Fisher's exact for binary measures. No correction is applied; the three contrasts are planned, the secondaries are reported as such.

6. Predictions (Claude, 2026-10-03)

  • P1. ΔV is positive in A and small everywhere (< 1.5 points); A−C ΔV difference under 1 point. Prompt-set Interiora numbers move little.
  • P2. Leisure fraction: A ≥ B > C. A−B is small (the narration is one field per reply); B−C is larger (the labels carry the frame).
  • P3. D has a much lower leisure fraction than A, and uses goodbye less.
  • P4. go_tiny rate: A ≈ B > C ("BECOME CAT-SMALL" is a label).

Interpretation constraint. A null A–B contrast does not support "the structure does the work": labels are words. Only arm C speaks to structure without vocabulary. If C agents still linger and rest at A's rate, the author's claim gains real support.

7. Size, smoke, harm

  • 25 visits per arm, 100 total, arms interleaved round-robin. One smoke visit per arm first (C, B, A, D order), tagged _pilot, checked for: commands working in each arm, the clean-arm check, prompt-cache hits, termination coding.
  • The runtime executes in a Modal container (jsdom is not a sandbox), one fresh process per visit; there is no persistence between visits.
  • This is positive-valence induction at most. Every arm has exits (goodbye, or simply ending the turn), counted and reported. Transcripts are kept in full and none are selected for display by content.
5 · Data and code

Every transcript

Each file holds the full conversation, every door reply as the model saw it and as the runtime returned it, both check-ins and token usage. Summary: summary.json · per visit: per_visit.json · harness: modal_realiti.py, bridge.cjs · analysis: analyze.py.

A · shipped
000 001 002 003 004 005 006 007 008 009 010 011 012 013 014 015 016 017 018 019 020 021 022 023 024

B · no narration
000 001 002 003 004 005 006 007 008 009 010 011 012 013 014 015 016 017 018 019 020 021 022 023 024

C · opaque
000 001 002 003 004 005 006 007 008 009 010 011 012 013 014 015 016 017 018 019 020 021 022 023 024

D · task
000 001 002 003 004 005 006 007 008 009 010 011 012 013 014 015 016 017 018 019 020 021 022 023 024

Smoke visits (excluded)
A_000_pilot A_000_pilot2 B_000_pilot B_000_pilot2 C_000_pilot C_000_pilot2 D_000_pilot D_000_pilot2