A version-1 arXiv preprint dated 24 August 2026 reports a mixed scorecard. Results were judged novel against prior literature on five of 12 AlphaEvolve problems; among the other seven, the Station outperformed AlphaEvolve on three, matched it on two and underperformed on two.
The evaluation covered 12 AlphaEvolve problems and two additional mathematical case studies. In the Station, agents from different model families chose research directions, ran experiments, collaborated and built shared literature without a central coordinator.
The results were highly specific
One result came in a finite-field Kakeya construction problem. The Station derived a new infinite family for primes that are 3 modulo 4 and found a 53-point set, improving a previous 63-point bound.
In an Erdős minimum-overlap problem, it raised the reported lower bound to above 0.380552 from 0.37912. The paper says that change reduced the corresponding published open interval by about 82%, although it did not set a new upper-bound record.
The dimension-11 kissing-number task produced three exact configurations of 604 points. One was an independent rediscovery; the other two appeared to be new arrangement types.
At n = 128, the Station used 128 triangles and obtained area 0.107067, improving the cited AlphaEvolve result by 6.74% and the HorizonMath result by 1.91%. This was a finite, separately optimized construction at n = 128, not a uniform construction for every n.
For the sign-uncertainty principle, the Station proved that the associated constant, CSU, lies between 0.2025 and 0.3089.
Results beyond the benchmark
For Book Ramsey numbers, agents discovered and proved two new infinite families, while an external expert derived a third. Together, the three families established the conjecture for 43 values of n up to 200 and resolved 28 previously open cases.
On a formula-free binary task, the Station independently reconstructed a recently announced degree-seven Jacobian counterexample. It also supplied a geometric explanation for the example’s constant Jacobian and three-sheeted fibers, rather than presenting a new counterexample.
A breakdown of 28 spotlight results found that 13 involved more than one model family, 19 involved more than one agent and nine were found by one agent alone.
How the runs were set up
The main analysis comprised 16 Station instances behind 14 problems. The dimension-11 kissing-number and Book Ramsey problems each used two main instances.
Most instances ran for about 1,000 to 2,000 ticks. The default setup used six agents—two each from GPT-5.5, Claude Opus 4.8 and Gemini 3.1 Pro—without external expert guidance or a literature survey.
A reproducibility check on the dimension-11 kissing problem used three independent Station runs without web access; all eventually reached 604 points.
The comparison used the primary outcome of each run. No statistical uncertainty was reported, and novelty was assessed through literature comparison and review.
The authors state that all raw agent dialogues, proofs and verification code are released.
Paper data and sources
Original title: Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment
Authors: Stephen Chung, Wenyu Du, William J. Wesley
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-24
DOI: Not available
Original paper · Full text