The arXiv preprint reports 60 medals for Praxist on a 75-task machine-learning benchmark, compared with 55 for Claude Code. In a separate fixed rocket simulation, its selected controller landed successfully on all 12,288 trajectories, while the starting artifact succeeded on 4.03% and Weco on 17.12%.
Those are local comparisons, not a general test of whether Praxist will outperform autonomous research systems. The benchmark used one sweep per arm, the rocket result came from a stated simulator protocol, and the trading headline used a policy selected after the fact.
Praxist's central mechanism is an artifact-to-lineage cycle. Reproducible artifacts become findings, findings are promoted to a frontier, the frontier shapes the next agenda, and final artifacts are reported with lineage records.
One benchmark sweep favored Praxist
The benchmark covered 75 Kaggle-derived competitions, split into 22 low-, 38 medium- and 15 high-complexity tasks. Praxist earned 60 medals, or 80.0% of the suite, including 49 gold medals. Claude Code with Claude Opus 4.8 earned 55 medals, or 73.3%, including 34 gold medals.
Recorded model spend was approximately US$3,054 for Praxist and US$38,370 for the Claude Code sweep. The study treated cost as resource context rather than using it to assign medals.
The comparison has narrow statistical footing. Each arm was a single local sweep, so it does not provide a repeated-sweep estimate of variation. The systems also used different base language models, and the benchmark records went through separate integrity-screening processes.
A perfect rocket score inside a fixed simulator
Under the frozen complete-validation protocol, the selected controller landed successfully on all 12,288 trajectories. The starting artifact succeeded on 4.03% and Weco, an external optimizer reference, on 17.12%. The validation also included 1,024 roll cases, and Weco evaluated 793 candidate controllers.
That result applies only to the stated setup: a low-order simulator with exact state feedback, fixed banks that were adaptively reused and a first-contact scoring rule. The evaluation did not test navigation, hardware, post-contact stability or the wider range of disturbances found outside the model.
The study does not provide an additive causal decomposition of the rocket gain. The reported improvement reflects cumulative changes along a parent chain, and the allocator was tested only after its immediate parent already had 100% success.
The trading result was selected after the fact
Over 28 walk-forward windows, the selected trading policy recorded a 1,864.5% cumulative return and a 53% calendar-time compound annual growth rate, or CAGR, compared with 23% for the paired baseline. The paper reports that as a 2.3-fold CAGR ratio.
The reported artifact used one seed over 29 evaluation cells, carried three hard constraint violations and was selected post hoc rather than as the campaign's promoted artifact. That selection means the result does not establish that the return would persist under a clean multi-seed evaluation or in live markets.
The robotics tests were less one-sided
In a LiDAR-inertial-visual SLAM test, COVSCHED had a mean absolute pose-error root-mean-square value, or APE RMSE, of 0.0501 metres across fourteen NTU-VIRAL sequences, versus 0.0937 metres for stock FAST-LIVO2. The evaluator also recorded a 72.4% average reduction in visual-path processing time.
Those raw error figures were not treated as an accuracy gain. When trajectories were re-associated under a common timestamp rule, the mean relative APE change was -0.09%, rather than the reported raw comparison of 45%. Differing timestamp conventions and pose associations confounded the original comparison, and the timing measure was not an end-to-end latency test.
The fusion benchmark used five scenarios, three random initializations per scenario and a 100-step horizon, allowing up to 1,500 surviving simulator steps per controller. Praxist recorded 1,264 surviving steps, compared with 1,222 for a reconstructed PCS-style controller, and had the lower common-horizon WNRMSE error score at the 95th percentile: 2.86 versus 2.99.
The ordering changed on the full horizon. The PCS-style controller had the lower WNRMSE p95, 4.42 versus 4.65, and completed 11 of 15 episodes, compared with 10 for Praxist. Both closed-loop controllers used privileged target values, the comparator was an in-simulator architectural reconstruction, and no controller met the official pass rule.
A synthetic scheduling test favored evidence debt
The scheduling evaluation was a closed synthetic experiment. Eight policies were run 1,000 times per scenario, for 4,096,000 policy runs across 512 scenarios; 348 scenarios met the physical-feasibility criterion and were retained for the reported comparison.
In the 348 physically feasible scenarios, the mature-evidence debt controller reached 99.85% quota success, compared with 98.65% for a Boolean maturity signal and 98.36% for a nonredundant thin token. The result is consistent with the value of evidence inheritance within this model, but the model simplified GPU and CPU use, failures, durations and research-plan inventories.
The evidence remains tied to its test settings
The evidence covers one local MLE-bench sweep, task-specific simulator or backtest studies and a synthetic scheduling design, with results tied to stated evaluators, fixed protocols, selected artifacts and particular model or hardware configurations. Several studies also involved post hoc selection or adapted reuse of evaluation data.
The study does not provide repeated-sweep uncertainty for MLE-bench or run-to-run variance estimates for deterministic SLAM.
It is an arXiv version 1 preprint dated 26 August 2026.
Paper data and sources
Original title: Praxist: From Experimental Artifacts to Solution Lineages
Authors: Jin Li, Ahmed Murtadha, Zhiyu Wang et al.
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-26
DOI: Not available
Original paper · Full text