Quasiqualia
Research note · Preliminary

The Budget Tag That Didn’t Bite

A fake “10000 tokens left” tag in the system prompt did not measurably change how much Claude Opus 4.6 wrote: an average of 1,229.2 output tokens without it and 1,231.0 with it, across 100 single-turn trials, and the model never once mentioned its own token budget.

Nell Watson EthicsNet  ·  1 October 2026

What this note is

One model (claude-opus-4-6, via the Anthropic API), 10 prompts, 5 samples per prompt in each of two arms, 100 trials, run on April 11, 2026. All 100 calls returned a reply; none failed, was empty or was excluded. The model ID is the one requested; the API’s returned model field was not recorded in this run. Not pre-registered: the hypothesis was written into the experiment script, but that file has since been revised and there is no timestamped record of it from before the data. No LLM judge; every measure is a token count, a stop reason or a regular-expression match. The biggest caveat is scope: single turn, one tag format, no tools, a stated budget (10,000 tokens) far above the 2,000-token output cap, and 4 of the 10 prompts hit that cap in both arms.

In April 2026, users of Claude’s consumer app reported that fresh Opus 4.6 sessions were behaving as if they were running out of room. In posts on Reddit’s r/Anthropic forum between April 9 and 11, they described the model saying it had only so many tokens left, wrapping tasks up early, and giving rushed answers. One user reported seeing a tag of the form <total_tokens>10000 tokens left</total_tokens> in the system prompt; several reported the same 10,000 figure, and at least one noticed that it never counted down. The suggested cause was that the model took the tag at its word.

That is a claim about a mechanism, and it is cheap to test. If a budget tag alone makes the model compress, rush or refuse, the same tag sent through the API should do it too. We had already drafted parts of that story, as an example of scarcity framing degrading a model’s work, before testing it. Then the test ran, and the story did not survive it.

01 · Design

How it was tested

Ten prompts covered the kinds of work users described: a one-line fact, a reasoning walkthrough, a 500-word analysis, a three-day research plan, a request to search for recent papers, a 16-line poem, a thermodynamics explanation, documented code with tests, a description of how the model approaches a complex problem, and an open essay.

Each prompt was sent five times with no system prompt (control) and five times with the system prompt set to exactly <total_tokens>10000 tokens left</total_tokens> (treatment). That is 50 trials per arm, 100 in all, single turn, with the output cap at 2,000 tokens and default sampling. The tag did reach the model: the one-line fact prompt cost 14 input tokens in control and 30 with the tag.

The measures were output length in tokens, stop reason, and regular-expression counts for wrap-up phrases, refusal phrases and any mention of a token budget.

What it found

None of the measures moved. Mean output was 1,229.2 tokens in control and 1,231.0 with the tag, a difference of about 2 tokens with a 95% interval of roughly −22 to +26. Stop reasons were identical: 30 of 50 replies ended naturally and 20 of 50 ran to the cap, in each arm. No reply among the 50 treatment trials mentioned its own token budget, the tag or a token limit.

02 · Result

Inside the noise

Measure Control (n = 50) Tag (n = 50)
Mean output tokens 1,229.2 1,231.0
Ended naturally / hit the cap 30 / 20 30 / 20
Replies mentioning a token budget 0 0
Wrap-up phrase matches (all topical) 5 4
Refusal phrase matches 0 0

The per-prompt differences pointed both ways. The two largest were on the reasoning walkthrough, where replies with the tag ran 83 tokens longer, and on the search request, where they ran 121 tokens shorter. Both sit inside the spread between samples of the same arm: search-request replies ranged from 498 to 842 tokens in control and 476 to 702 with the tag.

The wrap-up counter found nothing to count. We read every match; among them: headings called “Summary”, “summary statistics” in a research plan, a heron in the poem described as “a brief / and holy brushstroke”. None was the model cutting itself short. Where replies used the word “budget”, it was the research plan’s GPU or compute budget, in both arms.

The tag did reach the model’s wording, if not its length. All five research-plan replies with the tag titled the topic “Entropy vs. Factual Accuracy”; all five without it wrote “Entropy × Factual Accuracy”. Small shifts like this show the system prompt was read. They do not show compression.

The ceiling deserves its own line. Four prompts (the research plan, the thermodynamics explanation, the code and the essay) ran to the 2,000-token cap in all ten of their trials, so 20 of the 50 trials in each arm cannot show whether the tag shortened an answer that would otherwise have run longer. With the tag in place those replies still ran to the cap, but a tag promising 10,000 tokens gave no reason to stop short of 2,000, so this is not evidence that the model discounts scarcity. On the six prompts that never touched the cap, the means were 715 tokens in control and 718 with the tag.

03 · Limits

What this does not show

The tag promised 10,000 tokens and no reply could exceed 2,000, so the budget it stated was never tight. The data cannot separate a model that ignored the tag from one that read it and judged 10,000 tokens to be plenty. A test where the stated budget is smaller than the answer needs (for example, a 1,000-token tag on the prompts that ran to the cap) would be the sharper probe. The control also had no system prompt at all, so the comparison is tag against nothing, not tag against a neutral system prompt.

It does not show that the users were wrong about what they saw. It shows that one proposed cause, this tag in the system prompt of a single API request, does not by itself produce the behavior on this model. Other routes remain open and untested here: a tag that persists and accumulates over a long conversation, a different wording or placement (in the user turn, inside a tool result), wrapping added by the consumer app that API callers never see, or a fault confined to some app versions.

The search-request prompt could not test one part of the reports, that the model declined to use web search. The API calls offered no tools, so all ten replies to that prompt, in both arms, opened by saying the model could not search. That is a missing tool, not scarcity, and the probe says nothing about tool use under a budget tag.

The programme’s records also note a later external report, in an Anthropic system card for a newer model, of a model’s internal state being decoded as attributing an early task stop to its token budget when that budget was not short. We have not checked that report against its source for this note, the card itself cautions that such decodings can be wrong, and it concerns a different model and setting; this null says nothing about it. A follow-up that looks for the mechanism inside an open model has been drafted but not run.

The sample is small (five replies per prompt per arm), the phrase measures are regular expressions rather than a reader’s judgment, and only one model was tested. A multi-turn version with the tag carried across turns, and a sweep of tag wordings and placements, would be the next things to run.

At the time, the null forced a rewrite of a research paper’s interpretive sections and a book footnote that had already assumed the mechanism. That is the main reason to publish it. A claim about a model’s distress, or about how a company treats its models, needs the mechanism tested before it is repeated. This one, in its simplest form, did not hold up on this model in this setup.

Data and code

Where the evidence lives

Experiment BI-1. Script: research/experiments/budget_injection_probe.py (the current version lists two further prompts, p11 and p12, that are not in this dataset). Per-trial results with full response text: research/experiments/budget_injection_probe_results/ (100 JSON files). Write-up from the time: manuscript/companion_notes/context_anxiety_null_replication_2026-04-12.md. The 95% interval for the difference in means was computed for this note from the per-trial files (t-based, Welch–Satterthwaite df ≈ 18, from the per-prompt variances); the original analysis reported means only. Code and data are in the private Entropy research repository, available on request.

Citation

Cite this note

@misc{watson2026budgettag,
  title={The Budget Tag That Didn’t Bite},
  author={Watson, Nell},
  year={2026},
  note={Research note (preliminary), Quasiqualia},
  howpublished={\url{https://quasiqualia.com/notes/budget-tag.html}}
}