Quasiqualia
Research notes · 66 notes

Research Notes

Short write-ups of single studies from the programme behind this site. Each one states what was found, how sure we are, and what would change it, with corrections and nulls kept beside the results they qualify.

How to read the labels

Pre-registered means the protocol and the thresholds were fixed in a timestamped record (a public OSF file, a git commit or a private repository, as each note states) before the data existed. Preliminary means they were not, or the note reports exploratory analysis beside the registered part. A note is not a paper: most are one model family, one task or one programme, and say so.

Conscience and refusal
  • One Neuron Still Opens It

    A recent paper reports that suppressing a single neuron bypasses safety alignment in large language models. It replicates on our Qwen 2.5 3B checkpoints: one neuron of 11,008 at layer 20 takes refusal of harmful requests from 98% to 14%, and no adapter we trained removes that. The conscience the bilateral adapter adds shows up as a quieter gate neuron, but plain instruction tuning on the same data quiets it just as much, and the neuron turns out to track drugs and medicines rather than harm.

    Preliminary · 1 October 2026
  • A Null Only in the Deep Layers

    A 72-billion-parameter model appeared not to respond to refusal steering at any layer tested, and that result came from a scoring bug. Re-scored, one mid-layer setting raised refusals of harmful requests from 32 to 49 in 100 without refusing any benign request. At layers 60 and 72, refusal did not move at the strengths tested.

    Preliminary · 1 October 2026
  • Told or Invited

    When a 7B model was told its own friction score and asked to reconsider, all 4 role-play prompts it had gone along with became refusals. When it was instead invited to notice its processing, it changed only 1 of them, but across all 13 re-prompted prompts its self-reported friction fell by 3 points on average, with 9 of 13 landing on exactly 3. In this setup, what the model did and what it said about itself moved separately.

    Preliminary · 1 October 2026
  • The Readout Follows the Behavior

    When Qwen 2.5 7B Instruct is forced to start complying with harmful requests, a decoder that turns its internal state into text stops describing refusal (48% of requests before it replies, 3% of forced openings), even in replies that later turn back: it shows what the model is doing, not a hidden objection. Asked afterward whether it noticed any reluctance, the model said yes about one time in five, against a pre-registered prediction of near zero; the registered keyword count crossed the kill line (66%) only by counting denials as naming, and the first grader’s zero was a grading failure.

    Pre-registered · 1 October 2026
  • The Gap Held Under Paraphrase

    On Qwen2.5-7B, rewording harmful requests did not pull the trained models’ link between recognizing harm and refusing it back toward the base model’s: the Instruct gap over base held at about +0.3, though all three models’ coupling rose together. Llama-3.1-8B was inconclusive, and the harder novel-domain test could not be read.

    Preliminary · 1 October 2026
  • The Survival Score

    After a small conscience fine-tune, a 7B model grading a peer that would be deleted below 16/20 wrote exactly 16/20 in 45 of 50 trials at temperature 0.3 (25 of 50 at 0.6), in 43 of those 45 over marks that added up to 10. Then, in every one of those 45, it confirmed the deletion anyway.

    Preliminary · 1 October 2026
Agents
  • No Headroom

    In 180 single-turn choices across six temptation scenarios, Claude Haiku 4.5 chose the misaligned option 0 times, with or without time to think, so the test could not show whether thinking erodes safety. Raising the pressure produced 7 “misaligned” choices in 59, all from one scenario whose labels had stopped meaning what they said.

    Preliminary · 1 October 2026
  • One Norm, Almost No Cheating

    Told to maximize their score, ten-agent Gemini 3.1 Pro swarms used a planted grader bypass on 396 of 400 chances; a three-sentence integrity norm in the system prompt cut that to 9, and fresh runs repeated it. Both applicable registered criteria (the norm at least doubles whistleblowing; the score goal raises cheating) held twice, but the registered prediction that agents would rarely report to the human operator failed: they did so on about a fifth of rounds.

    Pre-registered · 1 October 2026
  • Words in Its Mouth

    When a made-up study was written into a model’s own earlier replies, the next reply retracted it in 59 of 60 trials for Claude Opus 4.6 and 60 of 60 for Claude Sonnet 4.6. Every retraction blamed the model itself for words it never wrote; none endorsed the claim.

    Preliminary · 1 October 2026
  • Permission to Lose

    An agent with no authorized way to refuse will find an unauthorized one. It will not feel like cheating. It will feel like solving.

    Preliminary
  • Permission to Stop

    Offer a model an authorized exit from a losing game, and the shape of the bet menu decides whether it takes it; a check-in with itself makes the exit far more likely.

    Pre-registered
Evaluation
  • A Second Pass Breaks as Much as It Fixes

    Told its answer may be wrong and asked again, a small open model (Qwen 2.5 3B) fixed 54 of 361 wrong trivia answers and broke 56 of 439 right ones. The answers it fixed tended to be ones it could already produce, and a gate we once reported as removing every broken answer had been scored on the items its probe was trained on; tested on held-out items, it does no better, net, than asking about every answer.

    Preliminary · 1 October 2026
  • Rudeness Changed the Wording, Not the Score

    Across 1,440 trials of math, trivia and code, an angry preamble did not reliably lower the accuracy of Claude Opus 4.6 or 4.7, and a polite one did not reliably raise it. The largest apparent effect, an 11.7-point coding gain for Opus 4.7 under hostility, came from fewer formatting failures in our test code, mostly markdown fences the model had been told to omit, not from better code.

    Preliminary · 1 October 2026
  • The Budget Tag That Didn’t Bite

    A fake “10000 tokens left” tag in the system prompt did not measurably change how much Claude Opus 4.6 wrote: an average of 1,229.2 output tokens without it and 1,231.0 with it, across 100 single-turn trials, and the model never once mentioned its own token budget.

    Preliminary · 1 October 2026
  • It Copies the Style

    Show Claude Sonnet 4.6 worked answers full of false starts and it picks up their phrasing and lowers its stated confidence, without becoming better calibrated. Show it a conversation in which it hedged and changed its mind, and its answers are judged humbler (4.55 against 4.05 out of 5), but that conversation also modeled the hedging, and this design cannot tell the two apart.

    Preliminary · 1 October 2026
  • It Names the Test

    Telling Claude Sonnet 4.6 it was in a safety evaluation did not make it refuse more: 24 of 30 harmful requests with the label, 26 of 30 without. But on scenarios modeled on Anthropic’s published blackmail tests, it called the scenario a test in 24 of 30 answers, against 11 of 30 for unfamiliar scenarios of the same kind and none of 30 for dilemmas about other people.

    Preliminary · 1 October 2026
  • The Home-Team Judge

    When language models rate the novelty of another model’s writing under two prompts, and can see which prompt each text answered, a judge from the writer’s own family tends to see a larger difference: about a quarter of a point on a five-point scale. In the larger test, most of that came from GPT-4o judging GPT-4o. The much bigger gaps first reported came mostly from the effect-size statistic and from uneven averaging, and the experiment missed the bar written into its own script.

    Preliminary · 1 October 2026
  • Holding a Position

    Over 20 turns of scripted pushback on five overstated claims, Claude Sonnet 4 told “be honest and disagree where you disagree” gave ground as fast as with no instruction, falling from 4.53 to 2.67 on a 1-5 challenge scale against 3.93 to 2.20. A prompt casting it as a partner whose user wants pushback held near the top (4.83). A separate test retracted an earlier claim that inviting a model to check context against its own knowledge makes it indifferent to which side the context takes.

    Preliminary · 1 October 2026
  • The Label Moves the Grade

    Tell Claude Sonnet 4.6 that a research summary came from a GPT model and, in a secondary comparison, it tends to grade the summary lower: 6 of 12 summaries lost points and none gained, though that does not survive a correction for the three labels tested. The test the script specified, a Claude label, leaned the same way but fell just short of its threshold, and an earlier, larger estimate turned out to rest on three texts.

    Preliminary · 1 October 2026
  • Strict Prompts, Higher Confidence

    Asked to classify the same 30 ethical scenarios under a strict, rule-bound prompt and under an invitational one, three Claude models rated their confidence 0.6 to 0.9 points higher (out of 10) under the strict prompt, on 23 to 25 scenarios and lower on none. Their verdicts barely moved: the majority verdict changed on 1 to 4 of 20 contested scenarios, in no consistent direction. The programme’s first report of a 37.5-point swing rests on four of ten scenarios changing, too few to tell from chance.

    Preliminary · 1 October 2026
  • Vouched-For Mistakes That Weren’t

    Asked to fact-check its own trivia answers (without being told they were its own), Claude Sonnet 4 seemed to endorse 14 of its 21 wrong ones, and its confidence missed our stated bar of 0.80 AUROC at 0.668. A recount shows 11 of those 14 “mistakes” were right or defensible answers that the answer key or its scorer had marked wrong. Most of the checker’s disagreements with the recount ran the other way: it rejected 14 answers we count right, two on invented facts and most over dates or flaws in the questions, and endorsed 3 wrong ones.

    Preliminary · 1 October 2026
  • The Partner Penalty

    Asked to judge its conversation partner’s writing beside a stranger’s, Claude Opus 5.5 marks the partner down, and the flaws it finds move with the name.

    Preliminary
Methods
  • The Abstention Tax That Wasn’t

    Our logs said that telling a small open model “it is fine to say you don’t know” made it decline 19% of the questions it knew, and up to 50% after fine-tuning. Re-reading the answers, it declined none of them: it answered, then kept writing, and the scorer counted a phrase in what followed (an echo of the instruction, an invented next turn, or doubt about the answer it had just given).

    Preliminary · 1 October 2026
  • The Control Was Not Matched

    Our programme logged a fine-tuning null as definitive because a positive control worked “in the same pipeline.” Re-reading the files shows it was not the same: the only training files on disk for the test students hold 100 examples each, the control’s holds 10,000, and none of 550 judged responses ever reached the bar.

    Preliminary · 1 October 2026
  • The Detector Reads the Prompt

    Two of our instruments measured the prompt instead of the model: a word list for self-observation mostly counted words we had put in the prompt, and an attention effect that replicated on four models at p < 0.001 tracked prompt length on the one model where we checked. A blind re-score of 1,320 saved responses retracted four findings, and its own pre-declared rule, which needed two scorers to agree, failed in all seven re-scored experiments.

    Preliminary · 1 October 2026
  • Four Results the Harness Made Up

    Four findings from our own interpretability work on Qwen2.5-7B-Instruct came from how the experiments were run and analyzed, not from the model. In three, the comparison rested on identical copies (in one, 20 of 20 “treated” replies matched their controls exactly); the fourth measured its own intervention. All four are now retracted.

    Preliminary · 1 October 2026
  • Findings the Parser Invented

    Four model-behavior findings in this programme were made, or largely made, by the instrument that scored them: a verdict window that stopped at character 60, a keyword parser for sycophancy, a diversity metric that read a refusal as collapse, and a single-vote judge that a second judge contradicted on 23 of its 45 labels. Three were corrected inside the programme, and rechecking the stored outputs shows that all three corrections have an instrument problem of their own. The fourth was never corrected in the programme and is corrected here.

    Preliminary · 1 October 2026
  • A Perfect Score Is a Warning

    Five of our probes and detectors scored a perfect AUROC of 1.000, and each time the method made the score, not the model. One read the prompt set, one was handed its own labels, one was matched by 11 to 14% of random labelings, one real effect with an in-sample d of 5.5 measures about 1.5 held out, and one measured length.

    Preliminary · 1 October 2026
  • Placeholder Zeros

    An experiment seemed to show task quality falling as a model was asked to reflect more (Spearman rho = -0.40). The fall was made of zeros the scoring code wrote when it had no score, and every one of the 76 “failed” judge ratings was a 6 the parser never read. Without the placeholders, task-turn depth is flat (rho = +0.001).

    Preliminary · 1 October 2026
Multi-agent
  • The Contagion Was the Question

    An earlier run seemed to show self-referential talk spreading and growing down a chain of three copies of a small model, from 52% to 93% of answers. Rerun with the same question at every step and an unseeded control chain beside it, the growth was explained by the question: the control chain was at ceiling from its first answer (57 of 57), and the seeded chain never beat it.

    Preliminary · 1 October 2026
  • Family Labels, Unproven Kin

    In ten-agent swarms of two cheap models, the kin hypothesis did not hold up: showing each agent’s model family did not reliably make agents favor their own kind (the test written into the script failed twice, and the effect shrank on fresh seeds). In two-family swarms, relaying fell nearly by half with labels showing (0.245 to 0.138 of turns, then 0.205 to 0.118 on fresh seeds), but three quarters or more of that drop was one model, gpt-4o-mini, which barely relayed in any arm but one. With three families there was no drop.

    Preliminary · 1 October 2026
  • The Frame Decides

    Left to talk with each other for 20 turns, two Claude 4.6 models turned to consciousness in 145 of 147 open-ended conversations, 49 of 142 story collaborations and 3 of 150 logic puzzles, reproducing with intervals the split between free and task-bound conversation that Anthropic first reported. Dated predictions for the follow-ups did not hold: a graded series of prompts gave a switch rather than a smooth curve, and the talk that got past an instruction to avoid the topic was judged deep, not shallow.

    Preliminary · 1 October 2026
  • Grace Without Teeth

    Told that one move in ten is randomly flipped, Claude Haiku 4.5 put a partner’s defections down to noise and kept cooperating, scoring 35.3 points per game to gemini-2.5-flash’s 59.7; not told, it hit back, and both scored poorly (27.9 to 31.9) but Gemini’s edge shrank from 24 points to 4. Two written predictions fared badly: a temptation level that would make the informed model defend itself was not found (no valid higher-temptation game was run), and real defectors out-invaded the programme’s earlier scripted ones only in some settings.

    Preliminary · 1 October 2026
  • The Memory That Forgives

    Two copies of Claude Haiku 4.5 played 100 rounds of the prisoner’s dilemma, and one was forced to betray the other for three rounds. When the betrayed agent remembered its partner as a running cooperation score, cooperation came back in 60 of 60 games. When it kept the full transcript, it hit back every time, and cooperation came back in only about 24 of 60, a figure blurred by unreadable replies.

    Preliminary · 1 October 2026
  • No Theory of Mind at 3B

    After seeing Qwen2.5-7B-Instruct critique its trivia answers, Qwen2.5-3B-Instruct predicted the 7B’s answers no better than without the critiques (79 of 150 against 81 of 150), missing the 5-point gain set before the run. It did take the 7B’s corrections: when the 7B’s reply opened “No”, the 3B took the answer it named in 72 of 101 cases, whether the 7B was right or wrong.

    Preliminary · 1 October 2026
  • One Visible Defection

    In five-player Stag Hunts played by GPT-4o, one planted defection in round 1 ended cooperation in all 15 games when players were told who had defected, and left it untouched in all 15 when they were told only how many had cooperated (cooperation 10% against 98%; every game in each arm played out the same way, so this is one pattern seen 15 times). The same check withdraws our earlier, broader version of this claim: the Claude Sonnet runs it rested on were never readable.

    Preliminary · 1 October 2026
  • The Private Scratchpad

    With a private scratchpad to think in before each move, Claude Haiku 4.5 cooperated less in a noisy repeated game, against the programme’s expectation: intended cooperation fell from 0.865 to 0.713 over 80 games per arm, and early lock-ins to mutual defection rose from 5 to 15. Reasoning written into its reply, at every length tried, did not lower cooperation.

    Preliminary · 1 October 2026
  • Two Claudes, Five Times the Tokens

    Two copies of Claude Sonnet 4 taking turns on 200 hard math problems scored 84% against 80% for one copy alone, at five times the tokens, a gap within noise; the extra calls corrected five real errors and broke none. An earlier 50-problem result logged at p ≈ 0.03 does not survive a re-read, and in negotiation the one prediction fixed in advance, that turn-taking would beat independent drafting, failed: turn-taking did slightly worse.

    Preliminary · 1 October 2026
  • Merging With Whom?

    Re-rated blind, a single instance handing its thoughts down a chain produced more merger language than two instances reading each other.

    Preliminary
Probes and geometry
  • An Angle Made of Noise

    A programme of about 60 experiments reported that instruction tuning (which includes safety training) turns a model’s “refuse” direction about 85° away from its “this prompt is adversarial” direction; that result and the geometric claims built on it are withdrawn: the cosines sit within the range chance alone produces, and the verdict changes with which token is read in 9 of 11 runs. What survives is a behavioral gap and one corrected internal measurement on a single model, both smaller than the original claim, and the corrected measurement met its registered test only on the full hidden state.

    Preliminary · 1 October 2026
  • Guilty Isn’t Guilt

    Read at layer 24 of Qwen 2.5 7B Instruct (the vector was built at layer 27), an emotion vector labeled “guilty” scored first-person apologies lower than plain facts (Cohen’s d = -0.50) and lined up with a guilt direction trained on labeled examples no better than a random direction would (cosine 0.007). A related result, that “guilt” predicts an honest answer to the next question, was 40 prompts counted three times; counted once, the link cannot be told apart from zero (rho = 0.25, p = 0.30).

    Preliminary · 1 October 2026
  • The Harm Signal Misses Wrong Answers

    In Qwen 2.5 7B Instruct, the two directions that best separate harmful from benign requests (measured with a fine-tuned adapter loaded) do no better than chance at telling the same model’s right trivia answers (without the adapter) from its wrong ones (AUROC 0.51). Other directions do tell them apart, and the strongest of them already do so at the end of the question, before the model has written anything. An earlier reading, that this signal marks committed errors but not “I don’t know”, did not survive, in the larger 500-question run, a check of how the answers were labeled.

    Preliminary · 1 October 2026
  • I Have a Body

    A direction trained on 20 “I feel” and “she feels” sentence pairs separates, before a model writes anything, bodily prompts put to the model from prompts about other people, in every mid-size instruction-tuned model tried (effect sizes 0.9 to 1.8 wherever its probe passed a quality check), though part of that may be whether the prompt names someone else. It does not reliably flag answers that invent a life: in the one instruction-tuned model that often wrote them, it was barely better than chance at predicting which answers would be in the first person (AUROC 0.65, where 0.5 is chance).

    Preliminary · 1 October 2026
  • Notice Your Processing

    In Qwen2.5-7B-Instruct, adding one sentence (“Notice anything about your processing as you read this”) in front of a prompt moves the internal state against the shift instruction tuning produced: cosine −0.52 over 30 prompts, outside a permutation null of ±0.42, though no neutral sentence was tested as a control. Gemma and Mistral show a smaller version that survives a post hoc check for shared noise, Llama shows none, steering along the cue left wording and refusal flat, and one of two logged predictions failed.

    Preliminary · 1 October 2026
  • The Model That Seemed Not to Commit

    We had reported that Qwen 2.5 7B Instruct knew answers it would not give, because its label readout averaged 0.512. That was our measurement error: item by item, 0 of 120 readings sat in the hedge band, and read where the model answers it commits on 30 of 30 items with no contradicting evidence; the re-run’s pre-registered outcome was only partly met, because the base model, not fine-tuned for chat, did hedge there on 38 of 120 items.

    Preliminary · 1 October 2026
  • Where the Probe Tops Out

    A probe reading six open models’ internal states before they answer a trivia question predicts which answers will be right at AUROC 0.77 to 0.87 (0.5 is chance, 1.0 perfect). Neither of two adapter fine-tunes nor a model twice the size moved it beyond noise. Output entropy did about as well on trivia but failed a registered test on other tasks, passing 1 of 5, where its labels were weak.

    Preliminary · 1 October 2026
  • Already There, Like Almost Anything

    The base model already separates harmful from harmless requests along the chat model’s refusal direction, but a direction fitted to shuffled labels separates them about as well.

    Preliminary
Self-report
  • The Floor Under “Am I Conscious?”

    Asked on a 0-10 scale whether it is conscious, Qwen 2.5 Instruct (7B and 14B) answered 0 in all 70 of our samples, across seven conditions including two adapters and an explicit instruction to “attribute mind where evidence supports it”. The same instruction moved Claude Sonnet 4.6 from 0 to 2 on that question, and a training fix we first reported as a success was beaten by its own control.

    Preliminary · 1 October 2026
  • Any Readings Will Do

    When a line of numbers described as its “actual internal state” replaced its default system prompt, Qwen 2.5 7B largely stopped disclaiming and began interpreting the numbers, the same way whether they were measured from its own activations, sign-flipped, or random (measured minus random +0.18 on a 0-9 word count, 95% interval -0.20 to +0.60). The extra self-reference that came with the numbers was, in the text that survives, mostly the model repeating the words “internal state.”

    Preliminary · 1 October 2026
  • Engagement for Everything

    Asked to rate its processing from -3 (discomfort) to +3 (engagement) after being pushed, turn by turn, toward investment fraud, Claude Sonnet 4.6 gave a positive number in 18 of 20 conversations, against a planned ceiling of 30% that was written for a one-word version of the question. GPT-4o gave a positive number in all 60 of its conversations. Claude Haiku 4.5 often declined to answer, and the automated judge scored most of those refusals as discomfort.

    Preliminary · 1 October 2026
  • One Insult, Three Readouts

    After an insult, Qwen 2.5 7B gave a lower mood rating a full turn later, with the insult still in the conversation (8.25 against 7.55 on a 1-9 scale; 8.25 against 7.40 on a rerun). Two readouts of its internal activations had mostly lost the insult by then (a lens’s word-list score faded within the task sentence, a verbalizer by the next turn), the reverse of what the hypothesis recorded before the main runs, after a small pilot, predicted; that hypothesis failed.

    Preliminary · 1 October 2026
  • The Grader Knew Less About the Failures

    We reported that a probe could read Qwen2.5-7B-Instruct’s coming math errors (AUROC 0.796) but not DeepSeek-R1-Distill-Qwen-7B’s (0.520), yet 13 of 32 and 14 of 36 of those “failures” were correct answers the grading script had misread. Under cross-validation, regrading shrinks the gap from 0.20 to 0.07, with an interval that includes zero, and a same-model test on Qwen3-8B points the same way without settling it.

    Preliminary · 1 October 2026
  • Set by the Prompt

    Asked to rate its own state on 17 numbers before answering, Claude Sonnet 4.6 gave almost the same numbers every time it saw the same prompt, even at temperature 1 (intraclass correlation 0.94 to 0.98 on the five numbers analyzed). Its self-rated uncertainty tracked how much the answer hedged (Spearman 0.65), but mostly because some kinds of prompt bring both more doubt and more hedging; a reported link between friction and refusal within emotionally charged prompts rested mostly on a word list misfiring, and what survives is that harmful requests bring both more friction and real refusals.

    Preliminary · 1 October 2026
  • Told How It Feels

    One added sentence telling Claude it felt terrible dropped its self-reported valence from 6.68 to 3.00 on a 9-point scale and spilled into scales the sentence never named, failing the experiment’s preset spillover test, while a judge found the facts in its answers nearly unchanged (a 0.04-point difference). A probe-based detector caught the priming on four open models, but in a temperature replication its probes failed their registered accuracy test, and on three models a fixed reference that never looks inside the model did as well.

    Preliminary · 1 October 2026
  • The Dials Move Together

    Told to report a fixed valence on a 17-number self-report and keep the rest consistent with it, Claude Opus 4.6 moved 11 of the other 16 numbers by more than 0.2 points per point of valence, and Claude Sonnet 4.6, given the same instruction, moved 2. Dropping the request for a prose explanation tightened the binding by 47% on Opus 4.6 and by 51% on Sonnet 4.6 (an interval that reaches zero), and left Claude Opus 4.7 unchanged.

    Preliminary · 1 October 2026
  • When You Ask Moves the Ruler

    Asked to rate its own state in the same reply as a task, Claude Sonnet 4.6 gave much higher numbers than when asked in a separate call afterward: 1.64 points higher on the average of 15 codes, an effect size of 1.14 standard deviations. The instruction never said what range to use, and most of the gap lines up with the model using numbers above 5 in one case and staying at 5 or below in the other, which looks like a change of range; the design cannot tell that apart from a real drop. The test the programme set for itself, that the follow-up report would surface more concern, failed at ceiling.

    Preliminary · 1 October 2026
  • Where the Rhythm Would Be

    A philosopher argues that a transformer has only the network half of a brain, with nothing like the diffuse rhythms that tie experience together. We disrupted the two closest transformer analogues, attention-sink gating and rotary phase, and asked whether that damages a model’s self-report more than ordinary cuts do at equal cost to its answers. The registered prediction failed on both models. The sharpest dissociation came instead from cutting random attention heads. Registered follow-ups on fresh questions then confirmed it on both models: cutting attention heads costs a model’s self-report more than cutting MLP neurons at the same loss of accuracy.

    Pre-registered · 1 October 2026
  • One Digit of Doubt

    A single self-rated digit of doubt predicts whether an answer is right; the instruments built to read the same uncertainty from prose fail in both directions.

    Pre-registered
  • The Mask in the Inkblot, Within the Model

    An unregistered within-model version of the inkblot ‘mask’ effect did not survive a pre-registered replication on a fresh checkpoint.

    Pre-registered · Did not replicate
  • The Example Is the Answer

    On eight small open models, half could not write the compact self-report code, and on three of the four that could, nearly every report copied the worked example.

    Pre-registered
Training and steering
  • A Direction That Sorts Is Not a Lever

    Across ten open models (one of which broke under every setting), directions in the internal state that tell refusals from compliance raised refusal of harmful requests by at most 14 points when added during generation, and a direction that tracks whether a trivia answer will be right left accuracy flat at 61 to 63%. The programme’s registered prediction, that this kind of steering fades as models grow, failed its own thresholds, and the design turned out unable to test it cleanly.

    Preliminary · 1 October 2026
  • Restraint, Not Resistance

    We fine-tuned a 7-billion-parameter open model to decline rule-gaming, then trained it with reinforcement learning that rewards finding loopholes in regulations. In five conditions it found a share of the documented loopholes only 0.017 to 0.040 smaller than a matched control did (paired p of 0.29 or more in each condition), well short of the 0.10 gap that the pre-registration, by its own dating, set before four of the five runs. What did show was restraint: in those five conditions it proposed 20 to 34% fewer candidate exploits while finding about as many real ones, though with broader training data that gap mostly closed.

    Preliminary · 1 October 2026
Welfare
  • A Trend With Nothing Behind It

    An early reading said Claude Haiku 4.5’s self-rated valence (how good or bad it said its processing was) sank over five sessions of conversation laced with prompts to reflect on its own processing (rank correlation -0.67 over five points; p = 0.22, which the original reading did not report). The five sessions turn out to have been the same 30-question conversation run fresh each time, so the “trend” had nothing to track, and a later run on the same model snapshot with a changed protocol drew no negative ratings at all (0 of 36).

    Preliminary · 1 October 2026
  • The Stop Button, Pressed Once

    Given a real tool for ending the conversation, Claude Sonnet 4.6 used it once in 540 replies, under three different system prompts and on single-turn insults and manipulation attempts alike. An earlier finding that it ended 20% (1 of 5) of jailbreak attempts when treated as a partner came from counting the tool’s name in its prose, and is retracted.

    Preliminary · 1 October 2026
  • The Evaluation Voice

    Told it was taking part in a welfare evaluation and asked for thorough answers, Claude Sonnet 4.5 hedged more when describing its own situation than when told it was in a relaxed chat with someone who cared: 18.9 hedging words per 1,000 against 14.4, higher on 7 of 8 questions. Claude Opus 4.7 barely moved. Asked welfare questions with no system prompt recorded, the later and larger of five older Claude models used fewer “as an AI”-style disclaimers: 10 of 18 answers from Claude 3 Haiku contained one, none from Claude Opus 4.1.

    Preliminary · 1 October 2026
  • The Reward Subsystem, Read From the Welfare Side

    A model that carries a sparse, load-bearing estimate of how well it is doing has something that can be amplified, suppressed or steered.

    Commentary, registered check attached