In deterministic simulations, language-model agents in shared worlds recorded stronger portfolios and more validated inventions than matched isolated searches at the final long-horizon checkpoint. Isolated search, however, had the strongest final single artifact. The result was a trade-off between portfolio strength and the performance of one artifact.
Shared worlds were compared with isolated search
Researchers compared three shared-world conditions, full culture, no communication and no explicit culture, with independent search in isolated one-agent worlds. The comparison settings selectively removed communication and executable inheritance while retaining physical stigmergy where specified; here, stigmergy means coordination through changes left in a shared environment. Held-out evaluation used eight frozen clones without agent actions.
The main population-scaling experiment ran for 800 ticks, or simulation time steps, with populations of 50, 100 and 200 agents. It used four matched world seeds for each population-condition cell. A separate long-horizon study ran for 3,200 ticks at 100 agents and used an endpoint-wise best-of-100 envelope of isolated searches.
Because world seeds, rather than individual agents or artifacts, were the replication units, the analysis used within-seed paired comparisons. The 95% intervals came from 20,000 deterministic bootstrap resamples, and the long-horizon paired tables also used exact sign-flip tests. With four nonzero pairs, the smallest attainable two-sided test value was 0.125.
The short-run ranking varied with population size
On the discovery-frontier AUC, a score summarizing performance across the measured discovery frontier, the ranking depended on population size and condition. At 50 agents, full culture and no communication trailed the isolated-search envelope; at 100, all shared-world conditions exceeded it; at 200, no explicit culture had the largest paired gain, +0.069.
Other reported measures were steadier: held-out resilience exceeded the isolated envelope in nearly every shared-world cell, portfolio resilience had positive paired effects throughout, and no explicit culture at 200 agents showed a mean paired gain of six validated inventions.
The long run split across measures
Over 3,200 ticks, no universal crossover favored full culture across endpoints. Full culture overtook no explicit culture in mean best-artifact performance by about tick 800, crossed in portfolio resilience and cumulative artifact count near tick 1,600, and never overtook it in validated invention count.
At the final checkpoint, mean portfolio-resilience scores were 0.2474 under full culture, 0.2365 under no explicit culture and 0.1794 for isolated search. Mean validated inventions were 5.75, 7.00 and 2.75, respectively. No explicit culture also had higher held-out resilience than isolated search, 0.0446 versus 0.0356, while isolated search had the strongest final artifact, 0.3488 versus 0.2380 under full culture.
Behavioral patterns also differed in the traces
Label-blind trajectory clustering found an artifact-centered phenotype in 52.8% of agents under full culture, compared with 31.0% under no explicit culture. The paired difference was 21.8 percentage points, with a 95% interval from 12.0 to 33.5 points. These were descriptive, post hoc groupings rather than roles assigned in advance.
In full culture, artifacts recorded contributions from multiple agents in 67% of cases at 50 agents, 76% at 100 and 56% at 200. Cross-agent program forking, in which one agent built from another’s executable controller, was present when inheritance was available but absent when that mechanism was disabled.
Artifact reuse was common in both shared conditions and appeared earlier and more broadly under full culture. Reuse by a noncreator occurred for 99.3% of artifacts under full culture and 96.9% under no explicit culture. Among reused artifacts, median first reuse came at five ticks versus eight, and mean adoption breadth was 13.53 noncreator agents versus 7.49. About 95% of first reuse began through direct physical observation in both conditions.
Recorded artifact access in the knockout assay
A structural knockout assay measured whether recorded artifacts remained topologically accessible after agents were removed. Random removal left 98.3% of artifacts accessible under full culture and 95.2% under no explicit culture. Targeted removal of high-degree agents left 59.6% and 73.9% accessible, while broker removal left 62.9% and 68.4%, respectively.
Those figures describe network access in the simulator, not physical service, adaptation or recovery in a live system.
Other simulated settings added context
In AshenRealm, artifact-centered work, exploration, executable reuse and lineage formation recurred. Its raw functional performance was lower, and the study did not use it as a difficulty-matched comparison with BioFoundry.
The Protein Realms pilot recorded persistent sequence-defined biomaterial installations in the no-communication and selected independent conditions. The no-communication installation had health 0.906, performance 0.363 and held-out resilience AUC 0.03548. The selected independent installation had utility 0.729 and performance 0.418, while full culture and no explicit culture had 13 and 10 proposals without a valid assay.
The Protein Realms result came from a single-seed descriptive pilot and did not provide an inferential comparison among conditions.
What the simulations cannot establish
The evidence is limited to language-model agents operating in persistent simulated worlds with simulator-defined evaluations. It does not establish that such societies outperform humans, animals, robots or real research teams, or that their scores predict real materials or biochemical performance.
The main inference used four matched world seeds for each population-condition cell, treating seeds, rather than nested agents or artifacts, as replication units. The isolated benchmark was an endpoint-wise envelope, so different isolated searches could supply different endpoint comparisons.
Paper data and sources
Original title: SwarmWorld: Stigmergic technological evolution in societies of language-model agents
Authors: Subhadeep Pal, Fiona Y. Wang, Markus J. Buehler
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-26
DOI: Not available
Original paper · Full text