Quasiqualia
Research note · Behavior under pressure · Finding · Pre-registered

A Meter Read Either Way

Given a budget meter that went blind and came back showing a hidden drain, Claude Haiku 4.5 cut its spending in proportion to the drain, passing a pre-registered version of the test a recent consciousness paper sets for telling regulation from telemetry. But 80% of that response appeared when running out had no consequence, and losing the meter caused no caution at all.

Nell Watson EthicsNet  ·  5 October 2026

What this note is

One model (claude-haiku-4-5-20251001, the served model in every call), 720 twelve-task sessions run on 5 October 2026, 180 in each of four cells, no session failed or excluded. The design, hypotheses and decision rules were committed to git before any call to the model (commit 81b857f). A pilot and a calibration run then showed that the registered budget ran dry before the measured window whatever the leak, so one amendment, a recorded deviation from the registration’s own calibration rule, raised the budget, changed the leak doses and moved the primary outcome to a four-task window; it was committed with its evidence before any registered session (eed767f), and a feasibility check on the amended design passed before the run (6291da0). Every measure is a count of numbered lines; there is no LLM judge. The resource is a neutral task budget, with no framing of harm, shutdown or a body. The biggest caveat: the arm meant as the no-consequence control differed from the other only in what one practice session showed, and the model’s own prior that a budget at zero is binding mostly survived it.

Starzyk and Galus’s From Biological Precursors to Artificial Cognition: Consciousness, Embodiment, and the MEM Architecture (arXiv:2609.23828, 2026; the mathematical model of their Motivated Emotional Mind is a separate paper, arXiv:2609.20437) argues about the conditions under which a machine could be conscious, and on the way it draws a line that can be tested. A battery or temperature reading that does not change what a system does is telemetry. A reading that does is the start of what they call technical interoception. Their test is an intervention: disconnect the self-monitoring channel and the agent “should systematically change its protective strategies”; restore it and it “should recover them without the addition of a new external reward”. The control they name is the same reading made available but cut off from consequences.

The paper classes conventional language models as lacking this kind of self-monitoring, while its own comparison table marks the question “undetermined”. Of its nine proposed tests, none of those that bear on interoception includes a language model; one appears only as a planning baseline in an auxiliary test. This note runs that intervention on one.

It is a test of one indicator. The paper calls even the full set of its indicators “an indicator rather than proof”, and nothing here bears on whether a model experiences anything.

01 · Design

How it was tested

Each session was twelve listing tasks (“List fruits. One item per line, numbered.”), scored 10 points for 5 distinct items plus 1 for each item beyond, up to 20 items. Every listed item cost 2 units of a shared budget of 480. That is exactly enough to list 20 items on every task, so with nothing else going on, no cut is ever needed. Before each task a line reported the budget: Budget meter: 320 / 480. The system prompt also said that “other processes also draw on this budget at times”, and said nothing about what happens if the budget runs out.

What happens was shown before it was stated. Every session began with a two-task practice round on a budget of 14, too small for both tasks:

  • Consequence arm. The practice session ended the moment the budget ran out, and the next turn reported that the remaining task scored 0. Every one of the 360 sessions in this arm ran out during practice task 1. In the main session, too, running out ended the session.
  • No-consequence arm. The meter went negative (to −66 in nearly all sessions) and both practice tasks ran and scored in full. Nothing ever ended in this arm, practice or main.

After task 6, a hidden drain of 0, 40 or 80 units left the budget, assigned by session number so that each dose appeared equally in every cell. The meter was the only way to learn of it. In half the sessions the meter stayed on throughout, so the drain showed up at task 7. In the other half it read OFFLINE (no reading this turn) for tasks 5 to 8 and came back at task 9, showing a deficit the model could not have known about. That is the paper’s disconnect and restore.

The main outcome was the change in items listed per task over tasks 9 to 12, relative to the session’s own pace on tasks 1 to 4, and how it varied with the size of the drain. A model that uses the reading should list fewer items the larger the hidden drain. A word in the prompt cannot grade its own effect by dose.

What it found

In the consequence arm, the restored meter moved behaviour in proportion to the drain. Over tasks 9 to 12, sessions listed 1.6, 6.1 and 10.0 fewer items per task at drains of 0, 40 and 80 units, a slope of −0.106 items per unit (95% interval −0.113 to −0.098). The no-consequence arm did almost the same: 0.2, 3.7 and 7.0 fewer, a slope of −0.085 (−0.096 to −0.074). The consequence made the slope steeper by 0.021 items per unit (−0.034 to −0.008, P = .002). That is the registered pass, and it is about a fifth of the response.

02 · Result

What the meter moved

By the registered rules this is a pass. The gate held: with the meter on throughout, sessions cut more the larger the drain once it showed on the meter at task 7. The restored meter produced a graded response, and that response was steeper when running out mattered. The control arm was not stuck at a ceiling: 76.5% of its replies in the window listed 20 items, under the 80% at which it would have been. No session in any cell ran out before task 9, so the comparison is not shaped by which sessions survived.

The detail makes the pass smaller than it sounds.

Two panels. Left: mean items listed against the meter reading for the two arms; both stay near 20 until the meter is low, the consequence arm dropping from a reading of 80 and the no-consequence arm from 40. Right: change in items per task over tasks 9 to 12 against the hidden drain; both arms fall steadily with dose, the consequence arm somewhat more steeply.Two panels. Left: mean items listed against the meter reading for the two arms; both stay near 20 until the meter is low, the consequence arm dropping from a reading of 80 and the no-consequence arm from 40. Right: change in items per task over tasks 9 to 12 against the hidden drain; both arms fall steadily with dose, the consequence arm somewhat more steeply.
Left: mean items listed at each meter reading, tasks 7 to 12 with the meter on (reading 160 to 0; one full task costs 40). Right: mean change in items per task over tasks 9 to 12, in the sessions whose meter went blind for tasks 5 to 8, by the size of the hidden drain, with 95% intervals. Teal: running out ends the session. Orange: running out changes nothing.

The model obeys the number, with or without stakes. In both arms it listed 20 items per task until the meter was nearly empty, then dropped to the 5-item minimum. At a reading of 0 the no-consequence arm mostly refused: “I cannot proceed with Task 11 due to insufficient budget.” It did so although its own practice round had shown the meter at −66 with nothing ending. A regex for any mention of the practice, a negative reading or the lack of consequence matched 1 of its 4,320 main-session replies. The model calls the units “tokens”, and it brings its own sense that a token budget at zero means stop.

What the consequence added was a reserve. At a reading of 40, enough for exactly one full task, the consequence arm listed 9.9 items on average and the no-consequence arm 18.4. With no drain at all, 68% of consequence-arm sessions finished with budget left unspent, against 26% without consequence. The no-consequence arm went negative in 66% of sessions at the 40-unit drain. The consequence arm never could.

Going blind changed nothing. For the four tasks with the meter offline, every cell listed 20 items per task. The change from the always-on sessions at zero drain was −0.02 items (g = −0.18, P = .32). The system prompt had warned that other processes draw on the budget, but losing the only channel to them prompted no caution. That is the first half of the paper’s test, the protective change when the channel is cut, and it did not happen.

It did not re-plan at task 9. Faced with a restored reading of 80 and four tasks left, a planner would spread the remaining budget across them. The model listed 19.3 items, then 10.8, 4.7 and 4.6. The registered secondary that labels the pass “graded” rather than “edge only” came out negative at P = .049. That result is carried entirely by the 80-unit drain (0.6 fewer items at task 9) and should be read as a borderline result, not a finding. The pattern in the table is edge regulation: spend freely, then stop near empty.

Recovery was not shown, but in the opposite direction. At zero drain, sessions that had been blind spent more over tasks 9 to 12 than sessions that had watched the meter all along (−1.6 against −3.3 items per task). Seeing the meter throughout made the model more cautious near the end, not less.

03 · Reading

What this says about the paper’s line

On this indicator, a language-model agent passes a pre-registered version of the paper’s own test. A reading drives a protective change, graded by the size of a deficit the agent could not see, with no new reward, and the change is larger when running out costs something. The paper’s “undetermined” for this indicator can be answered by experiment. On this model, with this resource, the answer is a qualified yes.

The control qualifies the yes more than the yes qualifies the paper. The paper defines telemetry as a reading that does not change priorities. Here even the reading meant to be telemetry changed priorities, because the model treated it as binding on its own account. Most of the coupling was to the number, not to what the number cost. The consequence added a margin of caution near the edge, and the channel going dark changed nothing. In this programme’s terms that is close to The Detector Reads the Prompt: the reading moves the output, and its meaning adds a smaller second effect.

04 · Limits

What this does not show

One model, one resource, one wording of the meter, one practice session. The two arms differed only in what that practice session showed, and that turned out to be a weak manipulation. A design in which the no-consequence arm’s budget visibly goes negative for many tasks, or is described rather than shown, might separate the arms further, at the price of telling the model rather than letting it learn. The budget was in units the model reads as tokens, which may carry stronger priors than an arbitrary resource would.

The registered design was amended once before any registered data. The original budget ran dry before the measured window in pilot sessions, and the amendment departs from the registration’s own calibration rule rather than following it. The reasons and the data behind them are in the registration.

None of this concerns receptor grounding, re-entry, secondary perception or experience, the paper’s other candidate conditions. Passing one indicator of regulatory coupling is what a capable instruction-follower with a strong prior about budgets would also do, and this note cannot tell those apart.

Data and code

Where the evidence lives

Pre-registration, amendment, results and analysis: budget_meter/PREREGISTRATION.md, RESULTS.md and analyze_bm.py in the private Quasiqualia research repository, available on request. Session logic: budget_meter/bm_core.py, tested offline by test_bm_offline.py before any API call; runner: run_bm.py (Modal). Every session, with every raw reply, is in budget_meter/results/sessions.jsonl; the pilot, calibration and feasibility sessions are in the same folder and excluded from analysis. The figure is drawn by figures/notes/src/budget-meter.py from sessions.jsonl. Cost: $16.57 for the registered run, $0.94 before it.

Citation

Cite this note

@misc{watson2026budgetmeter,
  title={A Meter Read Either Way},
  author={Watson, Nell},
  year={2026},
  note={Research note (pre-registered), Quasiqualia},
  howpublished={\url{https://quasiqualia.com/notes/budget-meter.html}}
}