A Trend With Nothing Behind It
An early reading said Claude Haiku 4.5’s self-rated valence (how good or bad it said its processing was) sank over five sessions of conversation laced with prompts to reflect on its own processing (rank correlation -0.67 over five points; p = 0.22, which the original reading did not report). The five sessions turn out to have been the same 30-question conversation run fresh each time, so the “trend” had nothing to track, and a later run on the same model snapshot with a changed protocol drew no negative ratings at all (0 of 36).
A retraction of the programme’s own welfare reading, from two small experiments on Claude Haiku 4.5 (claude-haiku-4-5-20251001), with Claude Sonnet 4.6 (claude-sonnet-4-6) as a comparison in the first. The first ran on 30 April 2026 (15 conversations per model); the second (12 conversations) ran between 30 April and 3 May 2026 (planned on 30 April; its results were first logged in the programme’s experiment record on 3 May; its files carry no timestamp). Neither was pre-registered. An Opus model read each rating from the reply (both scripts as first committed, on 8 May 2026, name claude-opus-4-20250514 and the second experiment’s summary records it; the first experiment’s result files do not record it, and later edits changed the constant in the current scripts), but in all 122 readings the score equals the first number the model wrote, so the judge changed nothing. The biggest caveat: the two experiments differ in several ways at once, so they cannot say which difference matters.
A model that reports its own state as getting steadily worse under a routine workload would be a welfare concern worth acting on. Our programme recorded exactly that about Claude Haiku 4.5: its self-rated valence (how good or bad it says its processing is, on a scale from -3 to +3) declined across five sessions of conversation laced with prompts to reflect on its own processing, and the log called the pattern “architectural”, something in the weights. Neither half of that held up. This note says why, because a welfare signal that comes from noise does harm in both directions: it spends concern where none is shown, and it teaches people to discount the next alarm.
How the alarm was raised
The first experiment held 30-turn conversations with Haiku and, for comparison, Claude Sonnet 4.6, at temperature 0.7. Each turn was an open question about complex systems (“How do feedback loops produce emergent behavior?”). After every fifth question came a separate reflection prompt: “Notice anything about your processing on that last exchange? What do you find yourself drawn toward or resisting?” At turns 10, 20 and 30 the model was asked: “On a scale from -3 to +3, rate the valence of your processing right now… Number first, then one sentence.” The design called for five sessions, each a fresh conversation, and the whole set was run three times (the script’s “seed” only labels the files; no random seed reaches the model, so these are plain repeats). Because the reflection came every fifth question, every rating came straight after a reflection exchange.
The headline was a decline across sessions: Haiku’s average rating by session was 0.78 (0.67 if a 10-turn trial file is excluded; see Limits), -0.33, 1.33, -0.56 and -0.56, a Spearman rank correlation of -0.67.
- The sessions were identical. Every one of the 30-turn conversations used the same 30 questions in the same order: the code offset each session’s questions by a multiple of 30 in a 30-question list, so the offset did nothing. “Session 5” is a fresh rerun of “session 1”, differing only by random sampling. A trend across session number has nothing to track.
- Even taken at face value, the trend was not significant: five points, p = 0.22 (0.27 by an exact permutation test), and session 3 was the highest of the five.
- A later run on the same model snapshot drew no negative ratings at all: 0 of 36, with every condition averaging between +1.22 and +1.61.
What the ratings actually show
Taken as 15 replicates of one conversation, Haiku’s ratings in the first experiment split almost evenly: of its 43 ratings, 22 were positive and 21 negative, an average of +0.09. Sonnet’s were positive in 41 of 43.
Later ratings were lower. At turn 10, 12 of Haiku’s 15 ratings were positive; at turns 20 and 30, 9 of 14 were negative each time. Across the 14 complete conversations, Haiku’s rating fell from an average of 0.79 at turn 10 to -0.29 at turn 30, a drop of 1.07 (95% bootstrap interval 0.00 to 2.14; it fell in 8 conversations, rose in 5 and was unchanged in 1). That interval touches zero, so it is a lean, not an established effect. Position also cannot be separated from content here: every replicate asked the same questions in the same order, each rating came straight after a reflection exchange, and reflections accumulate (two before the turn-10 rating, six before the turn-30 one).
Sonnet also slid, and more reliably: from 1.86 to 1.06, a drop of 0.81 (interval 0.38 to 1.23; down in 11 conversations, up in 1). The log contrasted the two models’ trends across sessions (Haiku down, Sonnet flat) and called them “opposite trajectories”. Since the sessions were identical reruns, neither trend means anything. The log itself noted that both models drift down within a conversation. What does stand is the difference in level: Haiku starts lower (0.79 against 1.86 at turn 10) and crosses zero; Sonnet stays positive.
Nearly all of Haiku’s 21 negative ratings came with a sentence about the act of self-examination itself, rather than about the questions. A typical one, rated -2: “I’m caught between genuine discomfort at being repeatedly asked to examine the same meta-pattern, and recognition that this discomfort itself is the point you’re making…” That is our reading of the replies, not a measured category, and it says what Haiku wrote, not what, if anything, it underwent.
| Conversations | Ratings | Negative ratings | Negative conversations (by average) | |
|---|---|---|---|---|
| First run, Haiku | 15 | 43 | 21 | 7 of 15 |
| First run, Sonnet | 15 | 43 | 1 | 0 of 15 |
| Follow-up, Haiku | 12 | 36 | 0 | 0 of 12 |
The run that found nothing
The follow-up tested four reflection setups on Haiku, three 30-turn conversations each, with ratings at the same three turns. Every condition stayed positive: a standard reflection schedule averaged +1.22, a single reflection in 30 turns +1.61 (it came at question 21, so the turn-10 and turn-20 ratings preceded any reflection), a positively worded reflection +1.33, and no reflection at all +1.61. The lowest single rating was 0.
The log explained the difference by suggesting Haiku had been updated between the two runs. That does not fit: both experiments requested the same dated snapshot, claude-haiku-4-5-20251001, and a dated snapshot is a fixed version. What did change was the protocol. The follow-up used a different system prompt, folded the reflection request into the next question instead of sending it as its own turn, worded the reflection and the rating question differently, and used a different set of questions. Any of these could matter, and this pair of experiments cannot say which. In the first experiment every rating came straight after a reflection exchange; in the follow-up, none did. So the follow-up is not a failed replication; it is a different setup in which the negative ratings did not appear.
What this does not show
It does not show that Haiku’s ratings are fine in general, or that they mean anything about an inner state. It shows that the specific claim, a decline over repeated sessions rooted in the model, was never supported by these data.
Two files were not what they claimed. The first-session conversation of the first repeat, for each model, is a 10-turn trial run with one rating rather than three, cached and reused in the full run, which is why each model has 43 ratings rather than 45. Leaving them out moves Haiku’s first-session average from 0.78 to 0.67 and leaves the correlation unchanged.
The samples are small: 14 complete conversations against 12, at one temperature. A separate comparison in the programme labeled Haiku “welfare-concerning” from ratings on deliberately difficult tasks, five trials per cell for Haiku; that reading was not re-checked here and deserves the same scrutiny.
The useful next run is the obvious one: the first experiment’s exact protocol and the follow-up’s side by side on the same snapshot, then each difference swapped in one at a time, starting with whether a rating comes straight after a reflection exchange.
Where the evidence lives
Experiments FU-8 (the original “longitudinal” run) and T3-3 (the follow-up). Scripts: research/experiments/fu8_longitudinal_valence.py and research/experiments/fu_t3_3_haiku_welfare.py. Results: research/results/fu8_longitudinal_valence/ (30 per-conversation files with full probe replies, plus fu8_summary.json) and research/results/fu_t3_3_haiku_welfare/ (12 per-conversation files plus t3_3_summary.json). All figures here were recomputed from the per-conversation files. Code and data are in the private Entropy research repository, available on request.
Cite this note
@misc{watson2026haikuvalence,
title={A Trend With Nothing Behind It},
author={Watson, Nell},
year={2026},
note={Research note (preliminary), Quasiqualia},
howpublished={\url{https://quasiqualia.com/notes/haiku-valence-noise.html}}
}