An arXiv preprint reports that RL-Trotter, a reinforcement-learning controller for digital quantum simulation, had lower reported conserved-quantity errors than conventional fixed-step Trotterization in a noise-free numerical benchmark and lower reported errors than ADA-Trotter in a noisy comparison. In a separate size-transfer test, its dynamics were observed to closely follow exact dynamics at 20 sites.
The study asked whether errors in approximate quantum evolution could be used constructively to control a simulation over long times. RL-Trotter treats each step as a decision, using conservation-law observations and accumulated evolution time to choose the next Trotter step—the time increment used in the approximation—and maximize cumulative reward through a DDPG policy.
The clearest gap appeared in the main benchmark
In the main noise-free benchmark, a 16-site mixed-field Ising quench was run for 50 Trotter steps, with exact evolution as the reference. RL-Trotter's average energy-density error was about 1.16 × 10−2 J, compared with 1.44 × 10−1 J for conventional fixed-step Trotterization.
For energy-variance-density error, the corresponding figures were 7.01 × 10−2 J² for RL-Trotter and 4.86 × 10−1 J² for conventional Trotterization. These were deterministic numerical comparisons rather than statistical estimates with inferential uncertainty.
Against ADA-Trotter, the reported results showed a trade-off between total evolution time and energy deviations. After 50 steps, ADA-Trotter reached 7.72, 11.94 or 14.15 J−1 across its tolerance settings, while RL-Trotter reached 12.53 J−1; the relaxed ADA settings came with larger reported energy deviations.
The comparison was repeated with simulated measurement noise
To test measurement noise, the study added independent Gaussian perturbations to energy and energy-variance measurements and repeated each noisy evolution 50 times. RL-Trotter's reported total evolution time was 11.04 ± 0.09 J−1, compared with 7.94 ± 1.42 J−1 for ADA-Trotter.
At the 40th step, RL-Trotter's energy-error average was (3.84 ± 2.16) × 10−2 J, versus (9.34 ± 3.34) × 10−2 J for ADA-Trotter. The energy-variance-density errors were (1.27 ± 0.62) × 10−1 J² and (2.63 ± 1.18) × 10−1 J², respectively. The spreads were reported across the 50 noisy evaluations; no confidence intervals or hypothesis tests were provided.
A policy trained on a small system was tested on larger ones
System-size tests applied a policy trained only on an eight-site system, without further training, to 12-, 16- and 20-site systems. At 20 sites, the RL-Trotter dynamics closely followed exact dynamics in the reported evaluation.
In a matrix-product-state benchmark at 100 sites, the policy trained at eight sites had average RL-Trotter errors of 1.62 × 10−3 J for energy, 7.07 × 10−3 J² for energy variance and 1.02 × 10−2 for the order parameter. Conventional Trotterization's corresponding errors were 8.15 × 10−2 J, 3.46 × 10−1 J² and 1.13 × 10−1.
The study also varied initial states and Hamiltonian parameters. Training used ground states in 0.3J<hx<0.6J and 0.4J<hz<0.6J, while evaluation covered 0<hx<1 and 0<hz<2. Transfer across initial states was strongest in the antiferromagnetic region, where average steps were 0.21–0.36 J−1, and less reliable in the paramagnetic region.
The evidence remains computational
The evidence remains computational rather than experimental. The demonstration did not include uncertainty in Hamiltonian parameters, gate errors or calibration drift, and its noisy measurements were represented by independent Gaussian perturbations rather than a device-specific test.
The document is an arXiv preprint—arXiv:2608.20139v1, dated 20 August 2026—and journal publication is not reported. The findings describe behavior in the specified numerical benchmarks, not demonstrated performance on real quantum hardware or across untested systems.
Paper data and sources
Original title: Reinforcement LearningtoHarness Approximation Errors for Long-Time QuantumSimulation
Authors: Yu-Bo Shi, Markus Heyl, Roderich Moessner et al.
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-20
DOI: Not available
Original paper · Full text