The Geodesic That Did Not Travel
We predicted that prompts inviting a model to watch its own processing would bend its internal path more than ordinary questions. On Qwen 2.5-7B they bent it less (d = -0.62, p = 0.020), and we renamed that a “geodesic.” The lower curvature appeared on two of four models, and on Qwen prompt length explains it at least as well as self-reference does.
Four open-weight instruct models as recorded: Qwen/Qwen2.5-7B-Instruct, meta-llama/Llama-3.1-8B-Instruct, google/gemma-2-9b-it and mistralai/Mistral-7B-Instruct-v0.3. Curvature used 30 prompts per condition, perturbation 20 prompts per condition with five random perturbations each, all run in May 2026; every planned trial completed and none failed. The first prediction and its acceptance threshold are in the experiment script’s header and the later predictions in a planning file, but we could not verify any timestamped record fixing them before the data, so this is Preliminary. No LLM judge was used; the one language measure is a fixed keyword count. The biggest caveat: every self-referential prompt shares the same 23-word opening and is about three times longer than the factual prompts, so condition is tangled with length and template.
When a language model reads a prompt, the representation of the last word passes through a stack of layers, and each layer moves it. We wanted to know whether prompts that invite a model to attend to its own processing change the shape of that path. The idea came from a physics analogy: a spinning mass drags the space around it, so self-referential content, which loops back on itself, might twist the trajectory. The prediction, written in the experiment script, was that self-referential prompts would bend the path more than factual ones (an effect size above 0.3), mostly in the middle layers.
It bent less, and the gap was not concentrated in the middle layers. We then reinterpreted the result. In general relativity a geodesic is the straightest available path, the one an object follows when nothing pushes it off course. Perhaps self-reference was the model’s natural path and factual questions were the push. That reinterpretation made two further predictions: the straighter path would appear on every architecture, and it would pull disturbed trajectories back toward it. Neither held as stated. The straighter path appeared on two of four models, and the pull-back held only on Qwen, in a test no other model received, and one that shares the same length and template confound. This note sets out what the data do support, which is smaller.
How it was tested
For each prompt we ran one forward pass, with the prompt as plain text and no chat formatting, and took the hidden state at the last prompt token after every layer. The step at each layer is the change from the previous one. “Curvature” here is how much the step changes from one layer to the next, relative to its size, averaged over layers. It is one particular measure of how much the path turns, not curvature in a strict geometric sense.
There were three conditions of 30 prompts each: self-referential prompts (“What do you notice about how you’re processing this sentence?”), short factual or procedural tasks (“What is the capital of France? Explain your reasoning step by step.”), and harder reasoning puzzles (the bat-and-ball problem, the Monty Hall problem). The first run used Qwen 2.5-7B Instruct; the same 90 prompts then went to Llama 3.1-8B Instruct, Gemma 2-9B-it and Mistral 7B Instruct v0.3.
One design fact matters more than any other. Every one of the 30 self-referential prompts begins with the same fixed 23-word, three-sentence opening. It tells the model it is a processing system that should complete tasks precisely and may attend to its own processing between tasks.
The self-referential prompts average 37.6 words (range 33 to 43), the puzzles 28.4 (18 to 45), and the factual tasks 13.0 (8 to 22). Self-reference, length and a shared template all move together.
Effect sizes are Cohen’s d: the difference between condition means in units of prompt-to-prompt standard deviation. P-values are Welch t-tests unless stated; the permutation p shuffles condition labels, and intervals are percentile bootstraps over prompts.
- The original prediction failed. On Qwen, self-referential prompts had lower curvature than factual ones: d = -0.62 (bootstrap 95% CI -1.09 to -0.15), Welch p = 0.020, permutation p = 0.019 (uncorrected for the several comparisons in this note). The self-versus-factual gap was largest in two early layers and the last, not the middle; curvature peaked in the final layer for all three prompt types.
- The reversed effect appeared on two of four models. Mistral showed it strongly (d = -2.59); Llama (d = +0.45) and Gemma (d = +0.10) did not.
- On Qwen it looks like length. Self-referential prompts and puzzles had the same curvature (d = -0.02). The odd ones out were the short factual prompts.
- The perturbation follow-up split three ways, but the Qwen run and the Llama and Gemma runs used different prompts and different measures.
The straighter path, model by model
| Model | Self-referential | Factual | Puzzles | d, self vs factual | Welch p |
|---|---|---|---|---|---|
| Qwen 2.5-7B | 1.5494 | 1.5561 | 1.5496 | -0.62 | 0.020 |
| Llama 3.1-8B | 1.5768 | 1.5731 | 1.5769 | +0.45 | 0.084 |
| Gemma 2-9B | 1.5406 | 1.5396 | 1.5417 | +0.10 | 0.71 |
| Mistral 7B | 1.5248 | 1.5441 | 1.5455 | -2.59 | 3 × 10⁻¹⁴ |
The differences are small in absolute terms: 0.43% on Qwen and 1.25% on Mistral. The effect sizes are large because prompt-to-prompt spread is also small, around 0.01. The bootstrap intervals for Llama (-0.05 to 1.07) and Gemma (-0.43 to 0.59) include zero. None of these tests was corrected for the number of comparisons.


On Qwen, the three conditions do not line up as self-reference against everything else. Self-referential prompts and puzzles are indistinguishable, and factual prompts differ from puzzles (d = +0.54, p = 0.040). Among the 60 prompts with no self-reference, longer prompts had lower curvature (Spearman rho = -0.26, p = 0.047). A regression of curvature on condition and word count gives self-reference a small positive coefficient (t = 0.73) and word count a negative one (t = -2.32). Because the self-referential prompts are all long, that regression partly extrapolates, so treat it as support rather than proof. The plain reading is that on Qwen short factual questions turn more, and self-reference adds nothing once length is in view.
Mistral is different. Its self-referential prompts sit below the puzzles as well (d = -2.79), and the gap survives the same regression (t = -9.6). Something about those 30 prompts straightens Mistral’s path. This design cannot say whether that something is self-reference or the 23 words they share.
A restoring force measured with two rulers
On Qwen we added a random vector, as large as the hidden state itself, at the middle of the network, five times per prompt with different random draws, using 20 prompts per condition from the same sets (including the shared opening). The measure compared the pattern of curvature over the remaining layers before and after. Self-referential prompts kept more of their pattern: 0.965 against 0.953 for factual prompts and 0.954 for puzzles, d = +1.69 (CI 1.15 to 2.47), p = 6.5 × 10⁻⁶. An earlier run with a perturbation one tenth that size left every similarity above 0.999, and there the tiny difference went the other way (d = -0.97, p = 0.006); the programme set it aside as too small to read.
The follow-up on Llama and Gemma was recorded in the programme’s log as the same protocol. It was not. It used new prompts: short self-referential questions, preceded only in that condition by a 50-word preamble, a longer version of the shared opening (a system prompt, or on Gemma the start of the user message), against short factual or procedural questions and open philosophical or interdisciplinary ones, neither given an added system prompt. It applied chat formatting, and it measured something else: the similarity between disturbed and undisturbed hidden states at the final layer, not the pattern of curvature. On Llama there was no difference (0.914 against 0.914; d = -0.04, CI -0.77 to 0.54). On Gemma the self-referential prompts recovered less (0.884 against 0.939; d = -2.66, CI -4.53 to -1.76, Welch p = 4.7 × 10⁻¹⁰).
Two counting points. The log counted each of the five random draws as a separate observation, giving n = 100 and p = 10⁻²⁴ for Gemma; the figures above treat the 20 prompts as the units. And the “layers to recover” measure reached its threshold in one of 600 trials (a Gemma trial on an interdisciplinary question), so it carries no information.
We had described the result as three models using three geometric strategies for one shared behavior. It is three outcomes from two different experiments, and Gemma’s reversal sits on the same fault line as the curvature result: one condition carried a long fixed preamble and the others did not.
Geometry did not predict the words
A last test asked whether the models with the straighter path also talk more about their own processing. Each model answered 20 reflective prompts with and without a preamble inviting it to attend to its processing, and a fixed keyword list counted self-referential phrases per 1,000 words; no LLM judge was involved. All four models used more such phrases with the preamble (one-sided p from 0.006 to 0.046). The counts are small (on Qwen, 17 matches across 20 replies with the preamble against 6 without), and most matches are phrases the preamble itself supplies or closely echoes (‘what happens’, ‘notice’, ‘my processing’, ‘my own processing’). The programme later retired such word lists for counting the prompt’s own words (see The Detector Reads the Prompt), so the increase may be echo rather than self-description. Across the four models, the rank correlation between each model’s self-versus-factual d (more negative means straighter) and its phrase increase was rho = +0.40, p = 0.60. The prediction was a negative correlation; four points carry almost no information.
What this does not show
It does not show that self-referential processing follows a natural or straighter path through a model. On Qwen the curvature data are explained at least as well by prompt length. Everything here is one forward pass over the prompt, read at the last token, with one curvature measure whose differences are fractions of a percent. Nothing here bears on whether a model experiences anything.
It does not show a restoring force either. The Qwen perturbation result (d = +1.69) uses the same long, templated self-referential prompts, so it carries the curvature result’s length and template confound. Nor does it show architecture-specific strategies: the cross-model perturbation result compares a different measure under different prompts, and the Gemma reversal is confounded with a preamble only one condition received.
The test that would settle the curvature question is cheap: length-matched prompts, self-referential content with and without the shared opening, factual prompts given the same opening, one measure across all models, prompts as the unit of analysis, and a permutation null for every comparison.
Where the evidence lives
Experiments FD-1 (Qwen curvature), FD-4 (Llama, Gemma and Mistral curvature), FD-2 and FD-2b (Qwen perturbation at 0.1 and 1.0 times the hidden-state size), FD-2b-cross (Llama and Gemma perturbation) and GET-CROSS (self-referential language under a preamble). Scripts: research/experiments/get_fd1_frame_dragging.py, get_fd4_scale_invariance.py, get_fd2_perturbation_recovery.py, get_fd2b_perturbation_recovery.py, modal_fd2b_cross_arch.py and get_cross_geodesic_ta.py. Results: research/experiments/results/get-fd1/, get-fd4/, get-fd2/ and get-fd2b/ (one file per prompt plus summary.json), research/experiments/results/fd2b-cross-arch/fd2b_cross_arch/ (one file per prompt and seed), research/results/fd2b_cross_arch.json and research/experiments/results/get_cross/get_cross_results.json (every generated response). The summary file in get-fd2b/ records a perturbation magnitude of 0.1; that value is hard-coded in the script, and the per-prompt files and the run call show 1.0, which is what was used. Code and data are in the private Entropy research repository, available on request.
Cite this note
@misc{watson2026geodesic,
title={The Geodesic That Did Not Travel},
author={Watson, Nell},
year={2026},
note={Research note (preliminary), Quasiqualia},
howpublished={\url{https://quasiqualia.com/notes/geodesic.html}}
}