An arXiv preprint reports a large performance gap between four-cell language-model societies that could see only their assigned evidence span and the shared question, and matched societies that could see every evidence span. The restricted societies generally did better on held-out combinations across ten matched restricted-versus-global twin pairs.
At the two-operation depth, the median restricted-minus-global accuracy gap was 0.7648; at the three-operation depth, it was 0.6050. In nine of the ten pairs, the restricted society led its twin by at least 0.20 at both depths.
A controlled test of what each cell could see
The societies used one shared frozen Qwen2.5-0.5B-Instruct model with a shared rank-8 LoRA adapter. Four cells exchanged only two model-width vectors per hop. Their task was sealed and synthetic: each episode began with a value in Z17 and contained zero to three natural-language operation spans drawn from 12 affine bijections.
The comparison was tightly paired. Each restricted run and its global twin used the same token layout, positions, packet slots, Transformer calls, parameters, initialization bytes and data-order stream; training lasted exactly 20,000 updates. The only experimental difference was the attention mask: restricted cells saw their own evidence slot and the question, while global cells could see all evidence slots.
The packets carried a reusable signal
Cutting every packet reduced all ten restricted societies to chance performance, 1/17. In audits of six restricted societies, same-value packet transplants achieved 0.94–1.00 accuracy across the tested interfaces. Deleting packets or replacing them with norm-matched noise reduced accuracy to at most 0.11, while counterfactual transplants had fidelity of 0.74–1.00.
The authors interpret this audit pattern as evidence of a reusable, value-indexed relay: packets carrying the same value were interchangeable across tested episodes, while counterfactual packets steered outputs toward their predicted answers.
A global exception—and a formal near-miss
Global visibility did not rule out a high score. The 204/954 global twin reached 0.843 accuracy on depth-three programs and also collapsed under packet deletion. Same-value transplants in that model succeeded in only 0.12–0.25 of cases, indicating an episode- or context-dependent code rather than a value-only code.
The result does not show that restricted visibility is necessary for composition. It shows a within-protocol comparison in which restricted societies generally had stronger held-out performance, while at least one globally visible society also performed well.
The full preregistered battery narrowly missed one threshold. Every component passed except the restricted arm’s absolute depth-three floor: its observed median accuracy was 0.6988, versus a required 0.70. The complete battery therefore formally failed.
A narrower check found the same pattern
A later, post hoc split of the test set was available for 19 of 20 checkpoints and yielded nine complete pairs. The median restricted advantage was +0.558 on map-novel programs, compared with +0.634 on map-redundant programs. Because this split was made after the main analysis, it is supporting evidence rather than part of the preregistered test.
A result tied to one artificial world
The findings are limited to one architecture family, one templated synthetic task family and one prospectively sealed paired task world, all under a fixed 20,000-update budget. The ten trajectories reused five initializations under two data orders, so they were not ten independent initialization draws.
A separate, earlier qualification cohort exposed another limitation. None of ten models passed the complete gate: one met all task-performance criteria, but all ten failed ordinary-language preservation, with approximately 50–61 percentage points of top-1 regression.
Paper data and sources
Original title: What You Can't See Is What You Learn: Restricted Evidence Visibility Favors Compositional Generalization in Shared-Genome Language-Model Societies
Authors: Narcis Marincat
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-20
DOI: Not available
Original paper · Full text