An autonomous coding agent nearly matched a benchmark in a simulated cell-edge wireless power-control task, scoring 1.4775 against 1.4850 for the reference, or 99.5% of the strongest reported benchmark. The same champion’s full-grid inference was measured at a fraction of the reference’s time, with the study reporting a speedup of roughly 600 times. The score was produced by the supplied simulator, and no measured wireless data were used.
The preprint asks a wider machine-learning question: whether the layer that designs a learned system can be handed over entirely to an autonomous agent. It tests that question in a simulated wireless setting, using a fixed held-out set across the evaluated tasks. The document is an arXiv version 1 preprint dated 26 August 2026.
The rules of the test
The network model contained seven wrapped hexagonal cells. It used COST231 path loss and Rayleigh fading, with evaluation on a fixed, pinned held-out set. In plain terms, the held-out set was a fixed set used to check performance. The result therefore describes a defined simulation setup, not a measured network.
The protocol separated an immutable evaluator, a sole mutable training script and a research charter. Each inner run was fixed at 2,000 Adam steps. The evaluator scored a fixed grid of 17 task pairings and set a 10-second inference budget for the entire grid.
Candidate changes were judged with pre-registered falsifiers. Repeated identical runs calibrated a ±0.0005 noise band, and a change was retained only when its held-out score exceeded that band.
From first design to champion
Over 26 hours, the campaign ran 81 unattended experiments across six architecture families. It began with a first working architecture that scored 92.2% of the reference. By the end, the campaign had closed 94% of the remaining gap.
The final score was close, but not identical, to the comparison point: 1.4775 versus 1.4850, equal to 99.5% of the strongest reported benchmark. This was a held-out score under the pinned evaluator and within the simulated setting.
On an Apple M2 Pro, full-grid inference took 2.52 seconds for the champion and 1,583 seconds for the reference. The paper reports that comparison as a speedup of roughly 600 times.
A result with a narrow perimeter
Another result concerns only Kq=1, the minimum-percentile setting. There, the model returned the max-min-optimal allocation for every value of the trained weights. A fixed 40-pass implementation reproduced that same allocation with relative error of 3.2 × 10−7 on held-out data. The exactness statement is limited to this minimum-percentile setting, not to every percentile target.
One parameter set served every network size and percentile target in the stated trained and evaluated range. That coverage is reported only within that range.
The boundary of the claim
That boundary is important: all reported scores came from the supplied simulator, with no measured data used. The study does not establish how the controller would perform on measured wireless channels or in a real deployment.
The campaign also represents one draw from a stochastic search. Its findings are budget-dependent, and other training or inference limits could favor other designs.
The authors state that relevant files, the complete experiment log, model weights and reproduction scripts were released. The document remains an arXiv version 1 preprint dated 26 August 2026.
Paper data and sources
Original title: Agentic Autoresearch for Cell-Edge Power Control: Radically Redefining the Researcher's Role
Authors: Ahmad Khan, Akram Bin Sediq, Sara Azadegi Naeini, Raviraj S. Adve
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-26
DOI: Not available
Original paper · Full text