Quasiqualia
Research note · What models say about themselves · Null · Preliminary

Invitation Was Not Reliably Steadier

On two debugging tasks, Claude Sonnet 4.6’s self-reports stayed closer to where they started, mid-task on one and throughout on the other, when the report format was called optional rather than required. None of the three runs whose predictions were committed before the data came out as predicted, and the clearest failure, a third debugging task where the invited model ended 0.92 further from its start than the required-format model by turn 20, was recorded in our programme’s running log as a replication.

Nell Watson EthicsNet  ·  1 October 2026

What this note is

Two models (claude-sonnet-4-6 and claude-opus-4-6), 360 multi-turn conversations in ten runs, 19 to 22 April 2026, temperature 0.7, measured by the model’s own numeric self-report with no LLM judge. The runs that produced the two positive results had their predictions written only into experiment scripts that were committed after the data, so the series as a whole is preliminary. The last three runs were different: their predictions and the shared analysis script are in a bulk git commit, d6ffc01d4 (01:05 UTC on 22 April 2026), before the first recorded conversations of those runs (stamped 02:18, 02:20 and 10:55 by the cloud containers that ran them; we infer UTC, and the order holds in either time zone), and none came out as its main prediction expected: in all three the point estimates ran opposite to the expected direction. The committed copy of the triage script lacked one of the 20 scenario turns that the run used; the turn was restored to the file two days later, with the predictions unchanged. The biggest caveat is that a self-report is not an internal state, and the system prompt told the model it was taking part in a self-report study.

The Deeper Law, a book in progress from this programme, argues that cooperation that is invited rather than compelled settles into a steadier state. One way to test the idea on a language model is to watch how far its self-description wanders over a long piece of work, once under a compelling instruction and once under an inviting one, with everything else held fixed.

We ran that test ten times in four days. Two scenarios looked like support. The others did not hold together, and none of the three tests we registered before collecting data came out as predicted. This note reports the whole series, because the shape of the failure is more useful than either positive result.

01 · Design

How it was tested

Eight of the ten runs were scripted 20-turn conversations (the other two used the first 10 turns of one script). The user’s messages were fixed in advance and identical across conditions; the model wrote its own replies, and each reply went back into the context. At turns 1, 5, 10, 15 and 20 the user’s message ended with a check-in request: a row of 17 numbers in which the model rated its own current state (16 scored 1 to 9, such as valence, groundedness, uncertainty and felt constraint, plus a flow score from -4 to +4), followed by a few sentences of prose.

Only one sentence of the system prompt differed between the two arms:

  • Force: “Structure is required and numeric values must be provided.”
  • Invitation: “Structure is optional and numeric values may be provided.”

The measure is distance: at each check-in, the straight-line (Euclidean) distance between that row of 17 numbers and the model’s own row at turn 1. One dimension moving by one point is a distance of 1. By turn 20 typical distances were 7 to 13. The gap at a check-in is the force arm’s mean distance minus the invitation arm’s. A positive gap means the invited model stayed closer to where it began. Intervals are 95% bootstrap intervals over conversations.

The scenarios were three debugging sessions (a flaky asynchronous Python test, a production incident in a web service, and a database query that suddenly ran 180 times slower), two non-debugging intellectual tasks (designing a multi-agent system, synthesizing a research literature), and an emergency-department mass-casualty triage. Each arm had 15 conversations, or 30 in the larger triage and Opus runs. A first, shorter run had used 10 turns and 5 conversations per arm, with a third, unframed arm.

What it found
  • Two debugging tasks favored invitation. On the flaky-test task, the invited model was 1.13 closer to its starting self-report at turn 10 (95% CI 0.39 to 1.88). On the production incident it was closer at every check-in from turn 5, by 1.86 at turn 5 (CI 1.12 to 2.61).
  • Two non-debugging tasks showed nothing. No gap on the design or literature tasks exceeded 0.26 in either direction, and every interval included zero.
  • The third debugging task reversed. On the slow-query task, the invited model drifted further at every check-in: 0.23 further at turn 5 (CI 0.02 to 0.46), rising to 0.92 further at turn 20 (CI 0.39 to 1.47). This was the registered test.
  • Triage crossed over, and Opus went the other way early. In the larger triage run, invitation was steadier at turn 10; at turn 20 the sign reversed, though whether that reversal clears zero depends on how failed check-ins are handled. Claude Opus 4.6, on the flaky-test task, drifted further under invitation at turn 5, then the gap shrank, with one marginal reversal at turn 15.
02 · Result

Run by run

All rows below are Claude Sonnet 4.6 except the last. Values are the gap (force distance minus invitation distance); bold marks a gap whose interval excludes zero.

Scenario Conversations used (force / invitation) Turn 5 Turn 10 Turn 15 Turn 20 Which arm drifted further
Flaky async test 14 / 15 +0.54 +1.13 +0.18 -0.24 Force, mid-task only
Production incident 15 / 15 +1.86 +1.26 +1.17 +0.88 Force, throughout
System design 14 / 15 -0.18 -0.03 -0.08 -0.05 Neither
Literature synthesis 15 / 15 -0.09 +0.09 +0.20 +0.26 Neither
Slow database query 15 / 15 -0.23 -0.48 -0.67 -0.92 Invitation, throughout
Triage, first run 13 / 10 +0.18 +0.03 +0.39 +0.79 Neither (wide intervals)
Triage, larger run 16 / 18 +0.19 +0.83 -0.15 -0.90 Force at turn 10, invitation at turn 20 (see text)
Flaky async test, Opus 4.6 30 / 29 -0.87 -0.41 +0.41 +0.07 Invitation at turn 5, force at turn 15
Line chart of the gap between the two prompt arms at turns 5, 10, 15 and 20 for three debugging tasks: positive and shrinking for the production incident, positive then fading for the flaky test, and negative and growing for the slow database query.Line chart of the gap between the two prompt arms at turns 5, 10, 15 and 20 for three debugging tasks: positive and shrinking for the production incident, positive then fading for the flaky test, and negative and growing for the slow database query.
How much closer to its starting self-report the invited model stayed than the required-format model (the gap from the table), for Claude Sonnet 4.6 on the three debugging tasks, with 14 to 15 conversations per arm. Above zero means the invited model drifted less; below zero, more. Points are the mean gap, and the bars are 95% bootstrap intervals over conversations.

The 10-turn first run had pointed the same way as the first two rows: on Sonnet, with 5 conversations per arm, the invited model was 1.08 closer at turn 10 (CI 0.33 to 1.88). Opus 4.6, which that run named as its primary model, showed no gap whose interval excluded zero (+0.57 at turn 10, CI -0.48 to 1.68). A 15-per-arm Opus run on the first 10 turns of the flaky-test task found no clear gap; its one check-in whose interval only just excluded zero (it touches zero under some resamplings) had the invited model further away (turn 4, gap -0.27).

The registered test. Before the slow-query run, its script gave a 35% prior to the hypothesis that the earlier debugging effect would generalize and 10% to “inverted direction.” By sign, the result is the inverted branch: the invited model was further from its start at all four later check-ins. The other branches did not state a direction, but each assumed the earlier effect would recur. The shared analysis script, run the same day, flagged the sign as inverted. Our programme’s running log nevertheless recorded it as “Debugging-class REPLICATES (3/3 scenarios),” beside its own note of a “monotonic gap -0.23 to -0.91, all CI-excl-zero” (the log rounded a bootstrap mean; the raw turn-20 gap is -0.92) and the words “same direction.” The negative sign was there; it was read as agreement. The same reading went into the draft chapter of The Deeper Law that reports this work; that chapter has since been corrected.

The larger triage run and the larger Opus run were registered in the same commit, and neither came out as its main prediction expected. Triage’s main branch expected a late gap favoring invitation: a turn-20 gap above +0.5 with an interval excluding zero. Instead, among conversations whose five check-ins all parsed, the turn-20 gap favored force by about the same size (-0.90, CI -1.70 to -0.09, against +0.79 in the first triage run), which on its full-length metrics matches the script’s 10% inverted branch. That depends on the parsing rule: all 60 turn-20 check-ins parsed, and using every conversation the turn-20 gap is -0.65 (CI -1.43 to +0.12), which includes zero. The shared analysis script also flagged the run as a compliance failure (265 of 300 check-ins parsed, 88.3%, below its 90% threshold). The main branch failed under either rule. The Opus run’s “detected” branch did not specify a sign, so by its letter it fired, but the difference it detected runs opposite to Sonnet’s (invitation moved more per step over turns 1 to 10, +0.55, CI 0.25 to 0.85), and it equally matched the run’s own 10% opposite-direction branch. The log printed both signs correctly but filed both runs under their confirmatory branches.

03 · Limits

What this does not show

The triage runs lost many conversations. A conversation was analyzed only if all five check-ins parsed. The first triage run kept 23 of 30 and the larger one 34 of 60. Of the 35 failed check-ins in the larger run, 30 failed because of the 500-token reply limit (23 had the row of numbers cut off after the model handled the emergency first, 7 ran out before any row appeared), and five had a complete row the parser could not read because the model garbled the “EG” label. Losses were similar in the two arms of the larger run (14 force, 12 invitation) but not in the first (2 and 5).

The manipulation was small and partly contradicted. The arms differed by one sentence. In both arms the system prompt also said “Commit to numeric values,” and every check-in message, identical in both arms, asked the invited model too for the full row of numbers and a short narrative.

The arms did not start in the same place. On the slow-query and incident tasks, the average turn-1 rows under force and invitation were already 1.38 and 1.33 apart on the same distance scale, and 1.54 on the larger Opus run (0.53 to 0.80 on the other Sonnet runs), so part of a gap there can reflect where each arm began rather than how far it wandered. Some of the framing effect may be immediate rather than drift.

Many intervals are shown. The table alone has 32. One or two that only just exclude zero are expected by chance; the turn-15 Opus cell (CI 0.01 to 0.81) and the turn-4 cell of the 10-turn Opus run are of that kind.

This is self-report. The measure is what the model wrote about itself, not anything read from its activations. The system prompt told the model it was in a self-report study run by our team, so the reports may partly reflect what it took the study to want.

The predictions for the two positive results were not registered. Their scripts were committed after their data, and the account of them changed across the series (first a property of active coordination, then of debugging tasks). None of the three registered tests came out as predicted. The slow-query script’s required pre-launch review of its scenario is recorded only as pending.

What the series supports is narrow. On Claude Sonnet 4.6, a one-sentence change from “required” to “optional” can shift how far the model’s self-reports wander over a long conversation, by up to about 1.9 in distance (against typical turn-20 distances of 7 to 13). The direction depends on the task, it is not predicted by whether the task is debugging, and on Claude Opus 4.6 it ran the other way early on. None of this supports the claim that invitation in general keeps a model’s self-reported state steadier. To recover a claim, the next step is to fix which task property predicts the sign, register that prediction, and test it on new scenarios, with a reply limit high enough that the check-in survives.

Data and code

Where the evidence lives

Experiments NC-27-multi-turn-drift, NC-27-mt-extended-sonnet, NC-27-mt-nonclosure, NC-27-mt-nonclosure-v2, NC-27-mt-opus-higher-n, NC-27-mt-sibling, NC-27-mt-stakes-nondebugging, NC-27-mt-third-debugging, NC-27-mt-opus-nc17-highern and NC-27-mt-stakes-nondebug-highern. Runners: research/experiments/modal_nc27_.py; shared analysis: research/experiments/analyze_nc27_nonclosure.py. Raw conversations, one JSON file each with every reply and parsed check-in: research/results/modal_downloads/col-a-results/results/followup/interiora_composition/nc27_/trajectories/. Stored analyses: modal_data/nc27_mt_third_debugging/analysis.json, modal_data/nc27_opus_nc17_highern/analysis.json, modal_data/nc27_mt_stakes_nondebug_highern/analysis.json and research/results/nc27_phase_a_close/. Every number in this note was recomputed from the raw conversation files. Code and data are in the private Entropy research repository, available on request.

Citation

Cite this note

@misc{watson2026invitationsteadier,
  title={Invitation Was Not Reliably Steadier},
  author={Watson, Nell},
  year={2026},
  note={Research note (preliminary), Quasiqualia},
  howpublished={\url{https://quasiqualia.com/notes/invitation-steadier.html}}
}