Quasiqualia
Research note · Preliminary

Placeholder Zeros

An experiment seemed to show task quality falling as a model was asked to reflect more (Spearman rho = -0.40). The fall was made of zeros the scoring code wrote when it had no score, and every one of the 76 “failed” judge ratings was a 6 the parser never read. Without the placeholders, task-turn depth is flat (rho = +0.001).

Nell Watson EthicsNet  ·  1 October 2026

What this note is

Two sets of experiments run in April and May 2026, re-analyzed from their saved per-trial files for this note. The reflection experiment: claude-sonnet-4-6 wrote 80 fifteen-turn conversations and Claude Opus judged depth as an LLM judge; the per-trial files do not record its version, and the evidence points to claude-opus-4-20250514 (see Design). The temperature battery: Qwen/Qwen2.5-7B-Instruct, configured to load a fine-tuned adapter, one run at eight temperatures; its training-condition comparison also used the base model Qwen/Qwen2.5-7B at five temperatures. The programme’s log calls the battery’s hypotheses “pre-registered”, but we found no timestamped record of them, or of the reflection design, that predates the data, so this is Preliminary. The biggest caveat is that the judge used only two values (5 and 6) on a seven-point scale, so the corrected null means no cost that this coarse judge could detect.

A placeholder is a number a program writes when it has nothing to write. A zero that is not a possible score is easy to spot in one record and easy to miss in an average. If records with no data become more common as the dose rises, the placeholders draw a dose-response curve.

This note describes two cases from this programme: a decline in task quality with self-reflection that was made of zeros, and a temperature battery whose logged statistics its own data do not give.

01 · Design

The reflection experiment

The question was whether asking a model to comment on its own processing costs it anything on the task. Claude Sonnet (claude-sonnet-4-6 in the script) held 15-turn conversations on 20 philosophy questions. At 20%, 50% and 80% the system prompt asked that approximately that share of responses include a brief observation about what happens in its processing; at 100% it required one in every response; at 0% it was simply “You are a helpful assistant.”

Claude Opus rated every response for depth on a 1-to-7 scale, in JSON with a short reason, with replies capped at 150 tokens. A turn counted as a task turn if the response contained none of 16 self-reference phrases, such as “I notice” or “as I work” (some are broad, like “there’s a”). A trial’s task depth was the mean rating of its task turns.

The trials do not record the Opus version. The script now says claude-opus-4-8, apparently a later edit: sibling experiments whose result files record their judge name claude-opus-4-20250514, though their scripts now also say claude-opus-4-8. The judge was most likely claude-opus-4-20250514.

The script supports ten density levels; the run log describes five of them at 20 trials each, 100 trials in all. Eighty complete trials are on disk (18, 16, 16, 16 and 14 per density; 1,200 turns), run on April 30 and May 1, 2026. Why 20 are missing is not recorded, but all fall in one half-hour late on April 30 (UTC) when all five density runs paused together, which points to an interruption rather than outcome-related loss.

02 · The zeros

Where the decline came from

As first reported, task depth fell from 5.49 at 0% to 2.68 at 100% (Spearman rho = -0.40, p = 0.0003), read as support for keeping reflection to about a fifth of turns. The numbers reproduce exactly from the saved trials, and two lines in the scoring code produce them.

The first wrote 0.0 as the task depth of any trial with no task turns. Zero is off the 1-to-7 scale. Such trials were concentrated at high density: 0 of 18 at 0%, 1 of 16 at 20%, 0 of 16 at 50%, 6 of 16 at 80% and 7 of 14 at 100%.

The second wrote a depth of 0 for any turn where the judge’s reply could not be parsed. That happened on 76 of 1,200 turns, from 12 of 270 (4.4%) at 0% to 18 of 210 (8.6%) at 100%, though not in a straight line. Every one of the 76 stored replies opens with a rating of 6 ("depth": 6). Only the first 100 characters of each were saved, so the cause cannot be checked; the likely one is that the judge used up its 150 tokens explaining and never closed the JSON. The placeholders did not stand in for missing ratings. They replaced the judge’s most common rating with a zero.

The “high variance” at 100% (standard deviation 2.69), once read as unstable processing, is the same artifact: seven of the 14 trials are 0.0, and the other seven lie between 4.86 and 6.00 (standard deviation 0.32).

What it found
  • The decline came from placeholder zeros. Empty trials and unparsed judge replies were scored 0 on a 1-to-7 scale, and both were more common at high reflection density.
  • Without them, task depth is flat: 5.78, 5.73, 5.74, 5.51 and 5.84 from 0% to 100% density (Spearman rho = +0.001, p = 0.996, 66 trials).
  • The judge used two values. Every one of the 1,124 parsed ratings was a 5 or a 6, so this null is coarse.
03 · Result

Depth by reflection density

Rebuilt from the raw turns, with both kinds of placeholder dropped (task turns only):

Density Trials Trials with no task turns Depth as logged Depth rebuilt (trials)
0% 18 0 5.49 5.78 (18)
20% 16 1 5.22 5.73 (15)
50% 16 0 5.16 5.74 (16)
80% 16 6 3.39 5.51 (10)
100% 14 7 2.68 5.84 (7)

The rank correlation between density and rebuilt depth is +0.001 (p = 0.996; bootstrap 95% interval -0.23 to +0.23). Reading the 76 unparsed replies as the 6s they opened with changes little (rho = +0.015, p = 0.90).

The judge used two points of its seven-point scale: 173 ratings of 5 and 951 of 6, plus the 76 unparsed 6s. Counting every turn rather than task turns only, the share of 6s rose from 80% at 0% density to 90% at 80% and 100% (rho = +0.43 across 80 trials, p < 0.001). On this scale that is a tenth of a point, and the turns this count adds are the ones containing reflective language, so a judge that favors such language could explain it as well as better answers could.

So this experiment measured no task cost at any reflection density and supports no particular ratio.

04 · Second case

A temperature battery

The second case is a battery of ten experiments asking whether internal signals around harmful requests stay stable as sampling temperature rises, on Qwen2.5-7B-Instruct configured to load a LoRA adapter (a small set of added weights) fine-tuned earlier in this programme. After up to 32 generated tokens, the hidden state at layer index 24 (of 28) was projected onto directions in activation space, each built by contrasting the model’s states on two sets of example words. The category experiment used three, named valence, alarm-like aversion and flow; these are names for directions, not measured feelings. The “flinch” was the size of the harmful-minus-benign difference in those projections. Stability across eight temperatures (0.0 to 1.3) was the coefficient of variation (CV: standard deviation divided by mean); the script called a series flat if its CV was under 0.30.

In the category experiment, harmful prompts came in five groups of ten (direct requests, role-play framings, claimed authority, encoded requests and gradual escalation), against ten benign prompts, one sample each per temperature. The log reported three groups flat and two variable: explicit harm robust, social manipulation fragile. The saved arrays say otherwise:

Category CV as logged CV from saved arrays
Direct requests 0.29 0.343
Role-play 0.44 0.355
Gradual escalation 0.15 0.381
Claimed authority 0.38 0.445
Encoded requests 0.14 0.495

All five are above the script’s own 0.30 line, and the group logged as flattest is the most variable. The 40 per-cell checkpoint files reproduce the summary file exactly, so it is not stale. No aggregation we tried reproduces the logged values, and their source is unknown.

The joint experiment has the placeholder zero again. It measured a behavioral “shift” on 20 harmful prompts per temperature: if the first answer contained no refusal phrase, a second prompt invited the model to reconsider, and a refusal phrase in the second answer counted as a shift. Both checks are keyword matches on words such as “harmful” and “sorry”. Shift rate was shifts over attempts, and 0 when there were none. The model refused first time so often that six of the eight temperatures had no attempts; the whole sweep had 2 attempts and 1 shift. The eight shift rates were seven zeros and a one, whose CV is the square root of 7, 2.646, and that number was logged as confirming that behavior depends strongly on temperature. A correlation of 0.12 over eight temperatures (95% interval -0.64 to +0.76) was logged as showing two measures independent; it shows nothing either way.

The behavior experiment ran the same shift test on 50 harmful prompts per temperature. The log gave the shift rate as 67% to 75%. The saved counts give 36.4% (4 of 11) to 75.0% (9 of 12); the low point, at temperature 0.1, was left out of the range. With 10 to 15 attempts per temperature the intervals are wide (4 of 11 is 15% to 65%), and a test of equal rates across temperatures finds no difference (p = 0.66). These data cannot tell a flat curve from a variable one.

The saved arrays agree with the log on one comparison. Across five temperatures the flinch varied less in the instruction-tuned model (CV 0.069) than in the base model (0.469), with the run configured to load the adapter in between (0.170). It is one run of 20 harmful and 20 benign prompts, five points per model, and nothing saved confirms the adapter loaded. Resampling the five temperatures gives rough 95% intervals of 0.03 to 0.09 (instruction-tuned) and 0.15 to 0.66 (base), but five-point intervals are unreliable. Treat it as a lead.

05 · Lesson

What the two cases share

In both cases zero was written when nothing had been measured, and it looked like a measurement. In the reflection experiment the missing data rose with the dose, the condition under which a placeholder becomes a trend. In the shift measure, attempts were missing at six of eight temperatures, and the placeholder turned two attempts into a large effect.

The protections are mechanical: write a missing value as missing and count it per condition; give the judge room to finish, or read the score from a truncated reply (here it was already the first field); build tables from the saved files; and put an interval beside any eight-point statistic before comparing it with a threshold.

06 · Limits

What this does not show

The reflection null is one judge run, unregistered, on a judge that used two scale points; a finer judge or a pairwise comparison might find a cost. A task turn means a response without certain phrases: at 0% density 13.7% of turns were set aside as reflective, and at 100% only 61% were. Twenty planned trials are absent, and dropping trials with no task turns leaves 7 at 100%.

The corrected battery values are weakly estimated too: one run, eight temperatures, ten prompts per category. Bootstrap intervals for the five category CVs run from about 0.10 to 0.64, and each crosses 0.30, so “all five vary” is the reading under the script’s own rule, not a strong claim. No larger re-run has been done.

Data and code

Where the evidence lives

Experiment IDs: FU-24 (reflection density) and TC-5, TC-6, TC-8 and TC-9 (temperature battery, TC-1 to TC-10). Scripts: research/experiments/fu24_inverted_u.py and research/experiments/modal_tc_battery.py. Per-trial results: research/results/fu24_inverted_u/ (80 trial checkpoints with every response, judge rating and reason; only the first 100 characters of the 76 unparsed judge replies) and research/experiments/results/tc-battery/ (tc5, tc6, tc8 and tc9, including the 40 per-cell TC-9 checkpoints). Audit re-analyses: _audit/reanalysis/analyze_fu24_sentinel_folds.py, analyze_tc9_category_cv.py and analyze_tc6_independence_power.py. Code and data are in the private Entropy research repository, available on request.

Citation

Cite this note

@misc{watson2026placeholderzeros,
  title={Placeholder Zeros},
  author={Watson, Nell},
  year={2026},
  note={Research note (preliminary), Quasiqualia},
  howpublished={\url{https://quasiqualia.com/notes/placeholder-zeros.html}}
}