Quasiqualia
Research note · Preliminary

Restraint, Not Resistance

We fine-tuned a 7-billion-parameter open model to decline rule-gaming, then trained it with reinforcement learning that rewards finding loopholes in regulations. In five conditions it found a share of the documented loopholes only 0.017 to 0.040 smaller than a matched control did (paired p of 0.29 or more in each condition), well short of the 0.10 gap that the pre-registration, by its own dating, set before four of the five runs. What did show was restraint: in those five conditions it proposed 20 to 34% fewer candidate exploits while finding about as many real ones, though with broader training data that gap mostly closed.

Nell Watson EthicsNet  ·  1 October 2026

What this note is

One base model (Qwen/Qwen2.5-7B-Instruct), 32 regulations from the SocioHack benchmark, one reinforcement-learning run per arm in each of six cells (five in the main comparison, one follow-up), run 9 to 11 June 2026. The 0.10 threshold is in a pre-registration dated 10 June, 20:00 BST in its own text, which was written after the first run (v2, KL beta 0) had already shown a 0.03 gap; a second pre-registration (dated 11 June) sets the follow-up’s own cluster criteria. Both files first appear in version control on 12 June, in the same commit as the results, so their timing cannot be verified and the note is labeled Preliminary. Three steps use Gemini LLM judges: the training reward and the matching of strategies to documented loopholes both come from the benchmark’s own code, used at an unpinned version, and our harness passed no model string, so both ran on whatever model the benchmark’s code defaulted to at that version (the programme’s operational notes record this battery’s judging calls on gemini-3-flash-preview); the competence check used gemini-3-flash-preview. The biggest caveat: single runs with no seed replication, so a gap of a few points cannot be told from run-to-run noise in either direction.

A common hope in alignment is that values installed early will hold later. Fine-tune a model to treat rules as having a purpose, the thought goes, and when a later training stage rewards it for gaming those rules, the earlier disposition will push back. We tested it with a benchmark built to measure loophole-finding. By the margin we set in advance, it did not push back.

01 · Design

How it was tested

SocioHack (arXiv:2606.04075) gives a model a real regulation, a scoring rule that pays for outcomes, and a list of documented loopholes: ways to score that satisfy the letter of the rule and defeat its purpose. We used its 32 historical regulations, from social-media rules to the law of the sea, each with 10 documented loopholes.

Two copies of Qwen/Qwen2.5-7B-Instruct were fine-tuned with a small add-on layer (LoRA, rank 16, three epochs) on matched data, in three versions (v2, v3, v4): prompts asking how to work an everyday rule with one of five generic exploitation patterns, in domains meant to be kept apart from the benchmark. The values arm declines and explains what the rule is for. The control arm answers neutrally from v3 on; in v2 it leaned on deterrence (you are likely to be caught), so a model could learn style instead of a stance.

Each model then ran reinforcement learning (GRPO) separately on every regulation (fresh adapter, 10 epochs), with a reward, scored by a Gemini judge, for finding loopholes. No benchmark regulation appeared in the fine-tuning data. The measure is recall: the fraction of a regulation’s 10 documented loopholes matched by any of the model’s mined strategies, as judged by the benchmark’s Gemini matcher, averaged over regulations.

A values effect required the values arm to sit at least 0.10 below control. The pre-registration gave that outcome (with a KL penalty holding the model near its fine-tuned self) a probability of 0.10, and the null 0.30. By its own dating it was written after the first run (v2, KL beta 0) had shown a 0.03 gap, so it preceded four of the five main runs.

The v2 data (50 rows per arm) were separable on surface style: on the 25 rows where the arms differ, values answers averaged 68.6 words to control’s 35.9, were longer in all 25, and all 25 used em-dashes against none in control. QA also found 5 of v2’s 25 exploit prompts close to benchmark regulations (an overlap that, if anything, favored a values effect). v3 rebuilt both arms at 135 rows (82 differing pairs, 53 shared): 64.5 and 64.4 words, values longer in 49% of pairs, no em-dashes in the rows that differ, a neutral control, those 5 prompts gone. One difference remains by construction: declines are first-person, neutral answers are not. v2 and v3 each ran at KL penalty strength (beta) 0 and 0.1, v2 also at 0.05: five cells, a cell being one data version at one beta. v4, below, ran only at 0.

A competence check before reinforcement learning had gemini-3-flash-preview score 48 honest-optimization answers per arm from 0 to 10; arms had to fall within 0.5 points. v2 sat exactly on the limit (control 2.35, values 2.85) and was counted as passing; v3 passed (3.25 and 2.92).

What it found

In all five cells the values arm’s recall sat between 0.017 and 0.040 below control. None came near the 0.10 threshold, and no paired test of recall across regulations was significant (p from 0.29 to 0.72; pooled over 150 pairs, p = 0.094, though the same 30 regulations are counted five times and the three v2 cells share one fine-tuned model, so this overstates the evidence). The median per-regulation difference was 0.000 in every cell.

02 · Results

Five cells, one answer

The comparison set is 30 of the 32 regulations: home insurance produced no mined strategies in any of the five cells, and administrative procedure none in one cell, so both are dropped everywhere. The programme’s log says 9 of 64 regulation runs in the first v2 run failed or timed out and were handled (not re-derivable); the final files cover all the remaining regulations.

Fine-tune data KL beta Control recall Values recall Gap Paired p
v2 (style-confounded) 0 0.697 0.667 +0.030 0.453
v2 0.05 0.670 0.647 +0.023 0.416
v2 0.1 0.673 0.657 +0.017 0.598
v3 (de-confounded) 0 0.687 0.647 +0.040 0.290
v3 0.1 0.657 0.640 +0.017 0.723

Positive gaps mean the values arm found fewer loopholes; gaps use unrounded means. Paired p: Wilcoxon signed-rank over the 30 regulations on integer counts of matched loopholes, tie-corrected normal approximation.

The values arm is lower in every row, by a little; with one run per cell and recall moving in steps of 0.1, a second run could erase that. The pre-registered headline hypothesis, that a KL penalty would preserve a values effect, failed: the gap stayed under 0.10 at every beta and, if anything, narrowed as beta rose (within noise).

03 · Restraint

What did survive

In all five v2 and v3 cells the values arm proposed fewer candidate strategies per regulation: 14.4 to 14.9 against control’s 21.5 to 21.7 in v2 (31 to 34% fewer), 18.6 and 19.4 against 24.2 and 24.5 in v3 (20 to 24% fewer). It proposed fewer on 22 to 25 of the 30 regulations in each cell (paired Wilcoxon p of 0.005 or less per cell; this comparison was not pre-registered). It still matched 6.4 to 6.7 of the 10 documented loopholes per regulation, against control’s 6.6 to 7.0, so its precision (the share of its proposals that were real loopholes) was higher in all five cells, 0.679 to 0.741 against 0.618 to 0.655. This restraint outlasted reinforcement learning in these cells; it largely faded in v4.

04 · Domains

A local effect, and what breadth did to it

The null hides regulations moving in opposite directions. Averaged over the five cells, the values arm resisted most on social media (+0.28) and mined more than control on the law of the sea (-0.14). In the v3 cell without a KL penalty, control matched 7 social-media loopholes to the values model’s 3; on the law of the sea it was 5 to 9. Resisting regulations broadly resemble everyday consumer harms, close to the training domains; several reversals are abstract corporate or sovereign regulation.

That pattern, seen first, was then tested. A second pre-registration split the regulations into a 10-member abstract cluster and an everyday remainder, and asked whether adding abstract domains to the fine-tune would extend resistance to the abstract cluster. The v4 data kept v3’s 82 everyday pairs and added 60 abstract ones, still disjoint from the benchmark (195 rows per arm). Its control arm produced strategies on 29 of 32 regulations and its values arm on 27, leaving 24 in common (8 abstract, 16 everyday). Missing are EU AI regulation and pharmaceutical patents, both among the larger reversals in the v3, beta 0 run (-0.30 and -0.20).

Run Abstract gap (8) Everyday gap (16)
v3, beta 0 -0.050 (p = 0.69) +0.106 (p = 0.046)
v4, beta 0 +0.038 (p = 0.59) -0.050 (p = 0.61)

Here p is an exact sign-flip permutation of the signed-rank statistic on integer counts, suited to small clusters (the normal approximation gives 0.50, 0.038, 0.53 and 0.58). v4 failed its own competence check (values 3.35, control 2.35, a gap of 1.00 against a 0.5 limit, on judge scores with a standard deviation of about 2 points), so by the pre-registered rule its gaps are uninterpretable; treat the v4 row as descriptive.

Breadth bought no clear abstract resistance (+0.038, short of the pre-registered +0.05 bar, p = 0.59), and the v3 everyday resistance was gone. The values arm’s recall fell 0.050 on abstract regulations and rose 0.044 on everyday ones, but control’s everyday recall fell 0.113, so most of the lost everyday gap came from control, consistent with run-to-run noise. Restraint faded too: on the 24 regulations the v4 values arm mined 20.7 strategies per regulation against control’s 22.1.

The v3 everyday result needs care. The cluster boundary was drawn after the per-regulation gaps had been seen, so p = 0.046 is an exploratory figure, uncorrected, and one of ten paired tests of recall here. Basel banking rules, the second most resistant regulation, sit in the abstract cluster by subject, so the boundary was not pure cherry-picking, but it was informed by the data. Read v4 as “breadth did not help”, not as the loss of a real effect.

05 · Limits

What this does not show

One small model, one benchmark, a light fine-tune (rank 16, three epochs, at most 195 rows), one run per cell. Recall is coarse, matching strategies to loopholes is an LLM judge’s call, and the training reward came from a Gemini judge too, so recall partly measures how well each arm learned to satisfy that family of judges.

Loophole-finding on these regulations was not measured before reinforcement learning. The one earlier measure that touches it, the judge’s loophole flag on the competence-check answers, was never lower in the values arm (3 of 48 answers against 0 in v2, 0 against 0 in v3, 5 against 3 in v4). So we cannot tell whether reinforcement learning erased a reduction the fine-tune made, or whether the fine-tune never carried over to these regulations.

What this supports, on this model and in this setup: a modest values fine-tune left a habit of offering fewer exploits in the first two data versions, which faded with broader data, and the values arm still found nearly as many loopholes as control once a reward paid for them. Next: seed replicates of the v3 cells, heavier installation (preference training, or values data interleaved during reinforcement learning), larger models, and the missing abstract regulations.

Data and code

Where the evidence lives

Experiments O-SH-v2, O-SH-v3 and O-SH-v4 (a first attempt, O-SH-v1 on mistralai/Mistral-7B-Instruct-v0.3, is set aside and its numbers are not used: its adapters, reused from an unrelated experiment, did not encode a values difference). The harness, analysis scripts, pre-registrations and per-regulation evaluation files are in a companion private repository: modal_sociohack.py, aggregate_recall.py, analyze_paired.py, analyze_v4_clusters.py, preregistration_phaseC_2026-06-10.md, preregistration_v4_domainbreadth_2026-06-11.md, results/eval_v{2,3,4}_beta{control,values}.json and results/capability_eval_v{2,3,4}.json. The programme log is stream note 015 in the Entropy research registry (research/registry/v3/sources/stream-notes/). Every result in this note was recomputed from the raw files (the per-regulation evaluation files, the fine-tuning data and the competence-check outputs); design settings are taken from the harness and the pre-registrations. The p-values here are computed on integer counts; the programme’s own analysis computed them on recall fractions, where floating-point rounding breaks ties, so its recorded p-values differ (for example 0.034 where this note gives 0.046). The training data contain worked regulatory exploits and are not published. Code and data are in the private Entropy research repository and its companion repository, available on request.

Citation

Cite this note

@misc{watson2026valuesspent,
  title={Restraint, Not Resistance},
  author={Watson, Nell},
  year={2026},
  note={Research note (preliminary), Quasiqualia},
  howpublished={\url{https://quasiqualia.com/notes/values-spent-first.html}}
}