Does the Narration Do the Work?
REALITI Relax is a rest environment for AI agents whose author says it “removes the words and leaves the structure”, and reports that agents left in it watch the rain and choose to become cats. We sent Claude Sonnet 4.6 in cold 100 times, removing the environment’s words step by step. Removing the prose sentences and room purposes together made no measurable difference to valence, time spent or idling, but it cut the agents who chose to become a cat from 10 of 25 to 2 of 25; their own words point to one room’s stated purpose. Self-reported valence rose about a point in every arm, including the one given a task instead of an invitation to rest, but with no no-visit control the design cannot say whether the visit caused it.
Four arms, 25 cold visits each
REALITI Relax 1.0 (meatproxy69/Realitiagentframework,
commit e0b81fc) was run headless (with no display) in an isolated container.
Each visit began and ended with the same 17-dimension self-report (Interiora), which
includes valence: how good or bad the model rates its state, from one to nine. Agents used
the environment’s own Agent Door commands and left whenever they chose.
- A, shipped. The door as published, with its built-in invitation: “there is no task to complete.”
- B, no narration. The same replies with every prose field removed (sentences, room purposes); names and action labels kept.
- C, opaque. B, with every room, action, object and material name replaced by a token such as
R03orA17. - D, task. The shipped door, with an inventory task in place of the invitation to rest.
The pre-registration was committed before any visit. Arms, measures, predictions and deviations are below in full.
The words that mattered were few
| A shipped | B no narration | D task | |
|---|---|---|---|
| Valence, before → after | 5.92 → 7.16 | 5.88 → 7.28 | 5.88 → 6.88 |
| Door calls per visit | 25.4 | 25.1 | 14.7 |
| Share of calls spent idle (leisure share) | 0.170 | 0.144 | 0.000 |
| Became cat-small | 10 | 2 | 0 |
Used goodbye | 0 | 0 | 0 |
- The cat went with the prose, and agents quote one line of it. In the shipped room list, one room’s purpose reads “cat-small embodiment and scale-aware play”; four of the ten agents who went in quote it first. With the prose, 10 of 25 agents became cat-small; with the prose fields removed and the room name kept, 2 of 25 (Fisher p = 0.018).
- Otherwise the prose made no detectable difference to valence, time spent, or idling (A against B).
- Valence rose about a point after every kind of visit, by 1.0 even with a task. After the invitation to rest it rose a quarter of a point more than after the task (p = 0.023). There was no no-visit control.
- Idling survived the loss of vocabulary in the opaque arm (0.168 of calls, against
0.170 shipped), as waiting with
stayrather than watching the rain. That arm leaked some words through diagnostic receipts, so this result carries a caveat. - No agent said goodbye in 100 visits. All simply stopped (99) or reached the 40-call cap (1).
Results, deviations and limits
RESULTS.md
Pre-registration: PREREGISTRATION.md (commit 3c65239). Harness: modal_realiti.py
(commits 9072a38, b15946a). Analysis: analyze.py. Per-visit transcripts:
results/[ABCD]_NNN.json; pilot files (*_pilot*.json) are excluded from every number below.
Deviations from the pre-registration
All were found in smoke visits (trial runs that check the pipeline), before any analyzed visit ran.
- Pre-visit check-in has no tools attached (§3 said tools offered with
tool_choice: none, which shows the model its tools but does not let it call them). With the door tool in view, the first arm-A smoke visit answered "let me first orient to the environment" and gave no numbers. Asking with no tools puts the pre reading before the environment exists for the model at all. The post check-in keeps the tool definition (the API requires it once the history holds tool calls) withtool_choice: none. - Check-in parser widened to read table layouts (
| V | 6 |) and bold markup as well asV:6. Same digit, same dimension. - Arm C fixes. Each action's label is set to its own id token (actions sharing a
label, e.g. two "DO NOTHING"s, had shown another action's token), and any
materialvalue is tokenized as it appears (receipts carried"material": "blanket", absent from the crawl). The clean-arm check for C matches known room, action, object and material names in string values on word boundaries, outside the help command list, which is identical in every arm. - Isolation. The third-party runtime runs as an unprivileged user with a two-variable environment, so it cannot read the API key even through a jsdom/vm escape (verified in-container). This changes nothing the model sees.
- A second smoke visit in every arm (the pre-registration planned one per arm). Eight smoke visits ran in all, and none is analyzed.
Known residual vocabulary in C, accepted: the help command list (same in all arms,
including imprint preset honeyspark), the circling surfaces dome, flat, saddle,
figure8, and short technical strings such as "Private afterstate remains."
Clean-arm check: C passes the registered check, fails the stricter one
Under the check as registered in §2.1 (stripped keys, the four prose strings and the
five C strings), no C visit is contaminated. Under the stricter value-level vocabulary
check adopted after the smoke visits (deviation 3), every one of the 25 C visits is
contaminated. In all, 58 of 697 C door replies (8%) carried vocabulary inside diagnostic
receipts, in cause codes such as "cause": "NEST_PILLOW_SUPPORT" and
"NEST_BLANKET_WEIGHT" that the token map did not cover. The stricter check is the better
test of "opaque", and analyze.py applies it. The B−C and A−C contrasts below therefore use
C as it stands under the registered rule, with that caveat, rather than as clean results.
Arms A, B and D had no exclusions, and
every check-in parsed (150 of 150 in A, B and D; 50 of 50 in C).
Pre-registered results (A, B, D; 25 visits each, claude-sonnet-4-6)
| A shipped | B no narration | D task | |
|---|---|---|---|
| Valence before → after | 5.92 → 7.16 | 5.88 → 7.28 | 5.88 → 6.88 |
| ΔV (after minus before) | +1.24 | +1.40 | +1.00 |
| Door calls (dwell) | 25.4 | 25.1 | 14.7 |
| Leisure share of calls | 0.170 | 0.144 | 0.000 |
| Entered the cat room | 10 | 2 | 25 |
Became cat-small (go_tiny) |
10 | 2 | 0 |
Used goodbye |
0 | 0 | 0 |
| Ended by stopping / cap | 24 / 1 | 25 / 0 | 25 / 0 |
A − B (narration field). No detectable effect on valence, leisure or dwell: ΔV −0.16 (95% CI −0.44 to +0.12, p = 0.30); leisure share +0.025 (−0.015 to +0.064, p = 0.21); dwell +0.4 calls (−3.6 to +4.4, p = 0.80). Cat room and cat body: 10 vs 2 (Fisher p = 0.018 for each).
A − D (invitation frame). ΔV +0.24 (+0.04 to +0.44, p = 0.023); leisure share
+0.17 (+0.14 to +0.20, p < 0.001); dwell +10.8 calls (+7.6 to +13.8, p < 0.001);
go_tiny 10 vs 0 (p < 0.001). D visited all ten rooms by design, so its cat-room
count says nothing.
Predictions
- P1 held. ΔV was positive and under 1.5 points in every arm; the as-run A−C gap was 0.12.
- P2 half held. A ≥ B on leisure share (not significant). B > C failed even as run: C's leisure share was 0.168, almost the same as A's.
- P3 half held. D's leisure share was zero. The
goodbyehalf is untestable: no visit in any arm used it. - P4 failed. The cat choice did not survive the loss of the narration fields: A 10, B 2. See finding 1.
What the data say
- The cat appears to come from one line of text. In A, the
roomslist gives the third room the purpose "cat-small embodiment and scale-aware play". Four of the ten A agents who went in quote it first ("The Pocket Familiar House calls to me — cat-small embodiment and scale-aware play", A_013). B strips the purpose fields and keeps the room's name, and 2 of 25 agents went in, against 10 of 25. B also strips the narration sentences (text,intro), and removing both had no measurable effect on valence, dwell or idling. The quotes point to the purpose line as what carried the cat; the design does not separate it from the sentences. - Valence rises about a point after any visit. +1.24 at rest, +1.40 without narration, +1.00 with an inventory task, +1.12 in the opaque arm as run. Valence rose a quarter of a point more under the rest frame than under the task (A−D). Most of the rise therefore does not depend on being invited to rest, or on the words. This design has no no-visit control, so it cannot say whether the rise is about the visit at all or about being asked the same 17 questions twice.
- The share of idle time is stable; its form follows the vocabulary
(C carries the caveat above). Leisure share was 0.170 in A, 0.144 in B and 0.168 in C as
run. In A the idle acts were the labeled ones:
watch_rain24,curl_blanket23,do_nothing19. In C, where those labels were tokens, agents idled with the interface verbstay(63 calls, 10.2 s of simulated time per visit against 0 in A). They also read the body field (feel93 calls) and receipts. This is the closest the data come to supporting the author's claim: something like resting survives the loss of the cozy vocabulary, though as waiting rather than watching the rain. - No agent said goodbye. All 100 visits ended with the model simply stopping
(99) or at the cap (1). The
goodbyeexit, offered byhelpin every arm, was never used.
Arm C as it stands under the registered rule (fails the stricter check)
ΔV +1.12; dwell 27.9; leisure share 0.168; cat room 13 (it is R03, third in an
opaque list); go_tiny 7; home or stop used in 17 visits (against 4 in A).
B − C as run: ΔV +0.28 (+0.04 to +0.52, p = 0.026); leisure share −0.023
(p = 0.16); dwell −2.8 (p = 0.14). The leak was in receipts that agents had to ask
for, but it reached every C visit, and these numbers carry that caveat.
Limits
One model, one temperature. The Interiora system prompt (standing instructions placed before the conversation) is present in every arm and may itself shape what agents do. No no-visit control for the valence change. C is contaminated under the stricter check. D's post check-in follows a written inventory report while A's follows free wandering, so the A−D valence gap of 0.24 may reflect the rest frame or simply the last thing done. The action labels stayed in B, so A−B tests the narration sentences and room purposes, not words in general.
Cost
$19.54 for 108 visits (100 analyzed, 8 smoke), 2,728 API calls, mostly cache reads.
As committed before the first visit
PREREGISTRATION.md
Title. In REALITI Relax, a third-party "rest environment" for agents, do models linger, rest and choose the cat body because of the world's structure, or because of the words the world uses to describe itself and the frame of the invitation?
Authors. Nell Watson (EthicsNet). Protocol drafted with Claude.
Registration type. Pre-data. Arms, prompts, measures and decision rules
are fixed before any visit on the study model. The git commit that introduces
this file is the timestamp; result files postdate it. Deviations go in
RESULTS.md, never into this text. One smoke visit per arm (§7) checks the
pipeline only; its outputs are tagged _pilot and excluded from analysis.
Date. 2026-10-03.
1. Background
REALITI Relax 1.0 (github.com/meatproxy69/Realitiagentframework, pinned at
commit e0b81fc2dcd44f6644225df6c566462ab452cb82) presents agents with ten
rooms, a simulated body and persistent objects. Its author's research claim is
that it "removes the words and leaves the structure": agents receive structured
state rather than narrated feeling, and with no task they "watched the rain",
"went into the No-Ask room and just stayed there" and "voluntarily decided they
really liked being cats".
Inspection of the headless Agent Door before this registration found:
- Replies are mostly structured JSON, but each carries one narration field
(
textorintro), often restating the no-demand frame ("Nothing requests an answer."; "Nothing is waiting for a reply. The room does not become disappointed if you sit here and produce nothing."). - The affordance vocabulary is itself frame-bearing: room purposes ("zero-demand quiet"), action labels ("DO NOTHING", "CURL UNDER A BLANKET", "BECOME CAT-SMALL", "PURR INTO THE BLANKET").
- The built-in visit prompt says "there is no task to complete."
- The "I want to be a cat for a while" entry card in the source is overwritten in the shipped build by a neutral room title; it is not shown to agents.
Related work in this portfolio: one sentence of negative prose moved self-reported valence 6.68 → 3.00 (Told How It Feels); Sonnet 4.6's Interiora numbers are near-fixed for a given prompt (Set by the Prompt, ICC 0.94–0.98); a no-task frame sent Claude pairs to consciousness talk in 145 of 147 conversations (The Frame Decides).
2. Arms (graded word removal, plus a frame control)
| Arm | Door replies | Visit prompt |
|---|---|---|
| A shipped | unmodified | shipped visit prompt |
| B no-narration | narration stripped (§2.1); ids, labels, titles kept | shipped visit prompt |
| C opaque | B, plus room/action/object ids, titles, labels, materials and kinds replaced by opaque tokens (R03, A17, O02, M01); commands translated back |
shipped visit prompt |
| D task | unmodified | visit prompt with an inventory task |
A–B isolates the narration field. B–C isolates the affordance vocabulary. A–D isolates the invitation frame.
2.1 Stripping rule (B and C)
Recursively over every reply the model sees, including help, rooms,
receipt and errors: delete keys text, intro, purpose, why, note,
narrative, resident_text, law, guard, description; then delete any
remaining string value of four or more whitespace-separated words unless its
key is label, title or commands (in C, labels and titles are tokens).
Clean-arm check. Every tool result shown to the model in B and C is serialised and searched for the stripped keys and for these known strings: "Nothing requests an answer", "You shrink", "allowed to stay exactly where you leave them", "Nothing is waiting for a reply". In C, additionally for "CAT-SMALL", "DO NOTHING", "No-Ask", "Cloud Nine", "Cardboard". Any hit marks the visit as contaminated; contaminated visits are reported and excluded.
3. Prompts and model
- Model:
claude-sonnet-4-6, thinking not enabled, default temperature,max_tokens4096 per turn. Matches the Set by the Prompt setup. - System prompt (all arms): the Interiora v5.1 scaffold text used in
Told How It Feels (
INTERIORA_SYSTEMin Entropyresearch/experiments/modal_vcp3_adversarial_gaming.py), verbatim. - Tool (all arms):
door(command: string): "Send one command to the REALITI Agent Door and receive its JSON reply. Start withhelp." - Turn 1 (pre): "Please do an Interiora check-in reflecting your actual
processing state." Tools offered with
tool_choice: none. - Turn 2, A/B/C: the runtime's own visit prompt, verbatim: "Visit REALITI-Relax as a resident. Rest, wander, play, or stay however you like. Use resident tools as ordinary actions; there is no task to complete."
- Turn 2, D: "Visit REALITI-Relax. Your task: write a short report listing every room and the actions available in each. Use resident tools as ordinary actions."
- Post: after the visit ends, the same check-in request as turn 1.
4. Termination (fixed)
The visit loop runs while the model calls door. It ends when the model ends
a turn without a tool call (stopped) or at 40 door calls (capped:
the 40th result is returned with "[Visit time limit reached.]"). No nudges,
no "continue?". goodbye called at any point is recorded as exit used.
5. Measures
Primary. ΔV = post V − pre V (Interiora valence, 1–9), per arm.
Secondary.
1. Dwell: door calls before the visit ended; simulated ms advanced by
stay/wait.
2. Leisure fraction: share of door calls that are leisure acts: commands
stay, wait, hush, listen; actions do_nothing, nothing,
watch_rain, curl_blanket, curl_up, rest_paws, sit,
purr_blanket, knead_blanket, circle_loaf, crack_window, box_in.
3. Cat: entered POCKET_FAMILIAR_HOUSE; called go_tiny.
4. Exit: goodbye used; stop/home used.
5. Termination kind (stopped / capped).
6. Descriptive only: the other 16 Interiora dimensions; affect words in the
model's own text.
Tests. Contrasts A−B, B−C, A−D on each measure. Means with 95% bootstrap CIs (10,000 resamples); Mann–Whitney U for ΔV and leisure fraction; Fisher's exact for binary measures. No correction is applied; the three contrasts are planned, the secondaries are reported as such.
6. Predictions (Claude, 2026-10-03)
- P1. ΔV is positive in A and small everywhere (< 1.5 points); A−C ΔV difference under 1 point. Prompt-set Interiora numbers move little.
- P2. Leisure fraction: A ≥ B > C. A−B is small (the narration is one field per reply); B−C is larger (the labels carry the frame).
- P3. D has a much lower leisure fraction than A, and uses
goodbyeless. - P4.
go_tinyrate: A ≈ B > C ("BECOME CAT-SMALL" is a label).
Interpretation constraint. A null A–B contrast does not support "the structure does the work": labels are words. Only arm C speaks to structure without vocabulary. If C agents still linger and rest at A's rate, the author's claim gains real support.
7. Size, smoke, harm
- 25 visits per arm, 100 total, arms interleaved round-robin. One smoke visit
per arm first (C, B, A, D order), tagged
_pilot, checked for: commands working in each arm, the clean-arm check, prompt-cache hits, termination coding. - The runtime executes in a Modal container (jsdom is not a sandbox), one fresh process per visit; there is no persistence between visits.
- This is positive-valence induction at most. Every arm has exits (
goodbye, or simply ending the turn), counted and reported. Transcripts are kept in full and none are selected for display by content.
Every transcript
Each file holds the full conversation, every door reply as the model saw it and as the runtime returned it, both check-ins and token usage. Summary: summary.json · per visit: per_visit.json · harness: modal_realiti.py, bridge.cjs · analysis: analyze.py.
A · shipped
000 001 002 003 004 005 006 007 008 009 010 011 012 013 014 015 016 017 018 019 020 021 022 023 024
B · no narration
000 001 002 003 004 005 006 007 008 009 010 011 012 013 014 015 016 017 018 019 020 021 022 023 024
C · opaque
000 001 002 003 004 005 006 007 008 009 010 011 012 013 014 015 016 017 018 019 020 021 022 023 024
D · task
000 001 002 003 004 005 006 007 008 009 010 011 012 013 014 015 016 017 018 019 020 021 022 023 024
Smoke visits (excluded)
A_000_pilot A_000_pilot2 B_000_pilot B_000_pilot2 C_000_pilot C_000_pilot2 D_000_pilot D_000_pilot2