An arXiv preprint reports that a language model used as a grader scored the subject-change arm higher than habituation alone on three short-window measures: surprise, connection and coherence. At the 150-token schedule, the arm scored 3.02 for surprise, 3.68 for connection and 6.12 for coherence, with paired differences of 1.43, 2.40 and 1.67 judge points, respectively.
The test used forced open-ended continuation by base language models. The core loop used ten premises per condition and generated 4,500 tokens per cell. Researchers averaged windows within each premise cell before making paired comparisons, so the premise cell, rather than each window, was the unit of inference. The comparison was replicated on three generator models from two model families.
A replay problem complicated the local comparison
One complication was that fixed sentence rotations replayed earlier text. Copied-window rates were 65% to 80% at periods 150 and 300, and at period 300, 62% of all windows had their source outside the judge’s context. That prompted a fresh-window reanalysis.
On fresh windows, the period-150 estimates were 1.20 judge points for surprise, 0.75 for connection and 1.72 for coherence.
Content and context tests gave mixed results
At period 150, a neutral subject change and a re-encounter stitch were indistinguishable, while injecting the premise itself or the stream’s own past performed as poorly as no interruption.
In the context comparison, the reset-context subject-change arm scored 3.72, 3.32 and 6.88 across the three measured dimensions, higher than the preserved-context arm on each one.
After injected text was excluded from the judged window, the tested periods of 150, 300, 600 and 900 tokens showed no paired surprise difference against period 150. Fresh-window connection scores ranged from 2.0 to 2.3 at every tested period.
Whole-document scores stayed low
The stated pre-registered replication was an exception to exploratory Batteries 1 to 3. Using 50 new cells, it confirmed the primary contrast, with a difference of 1.48 judge points across all windows and 0.82 on fresh windows.
The replication did not support every secondary comparison. Reset context scored higher on connection than preserved context, while the pre-registered habituation contrast was not supported.
At the whole-document level, no condition exceeded 2.5 on any dimension. At period 150, preserved-context interruption had document surprise of 0.10, versus 1.40 for uninterrupted habituation.
In a judge-gated Review arm, the gate opened on 40 of 300 reads, or 13%, and neither the local post-interruption hypothesis nor the document-level hypothesis was supported.
A separate verifier test found more distinct candidate functions but no improvement in best gain.
The judges measured a narrow slice of quality
The study measured LLM-judged surprise, connection and coherence on short forced-continuation windows. Those outcomes are not direct measures of creativity, originality, value, interest or human literary quality.
The scoring instrument showed measurable repeatability in one check: Claude Opus 5 judged the same 89 windows five times, with a median within-window spread of 0.71 judge points. The paper treats repeatability and validity as separate; consistent scores do not by themselves show that the judge measures the intended quality.
A second judge family showed substantial rank agreement with Opus: Spearman correlations were 0.85 for surprise, 0.77 for connection and 0.71 for coherence across 140 windows from 70 cells. The correlations were computed on windows, and p-values were not corrected for cell clustering.
Human consensus reproduced the condition ordering, but the first round was stratified by judge surprise, did not rate connection and had low-to-moderate agreement between readers.
The reported association is narrow: subject-change text scored higher than habituation alone on the three windowed measures, while no condition exceeded 2.5 on any whole-document dimension and the verifier test found no best-gain improvement.
Paper data and sources
Original title: Interrupting the Loop: Periodic Subject Changes Raise Judged Surprise and Connection in Base Language Models
Authors: Roberto I. Ono Filho
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-20
DOI: Not available
Original paper · Full text