Preprint

LM-X posts higher robot success, but signals stay uncalibrated

Preprint: LM-X combines robot actions with online signals for progress, event intention and local action reliability.

LM-X, a robot-control policy that predicts task progress, the next meaningful event and local action reliability while acting, recorded higher average success than an action-only backbone in a pretraining gate and than GR00T N1.7 in both simulation and physical-robot comparisons. The results are benchmark averages: they show how the tested systems compared, not that the extra signals caused the gains.

On the 50-task randomized-hard RoboTwin2.0 benchmark, LM-X posted 74.1% mean success compared with 55.4% for GR00T N1.7, an 18.7 percentage-point difference. In the real-world evaluation, it recorded 68.6% versus 50.7%, a 17.9-point difference. The reported representative-task comparison in simulation and the real-world comparison were not uniformly favorable across tasks.

What LM-X adds

The study asks whether a generalist vision-language-action policy can learn three predictive states together, show them online and use them inside control. Those states are return-to-go, or RTG, for progress; event-to-go, or ETG, for the next semantic event; and propagated flow variance for local action uncertainty. LM-X is directly supervised to produce all three online outputs, and they participate in the action pathway.

The pretraining mixture contains more than 20,000 hours of real-robot trajectories, including over 1,000 hours of failed policy rollouts.

The early gate favored the full combination

Before the broader benchmark, the variants were trained from scratch on 45 RoboTwin2.0 tasks, while five disjoint tasks were held out for the pretraining gate and downstream post-training and evaluation.

On that five-task gate, LM-X had 79.6% mean success versus 63.6% for the action-only backbone, a 16.0-point difference. It was also 10.8 points above the strongest single-component variant.

The ablation results also showed task-specific brittleness in isolated auxiliary objectives. On the open-microwave task, the event-only variant scored 16% success, compared with 54% for the action-only backbone.

Higher averages across two test settings

The full randomized-hard RoboTwin2.0 benchmark used 50 randomized-hard demonstrations per task and 100 trials per task, with all 50 tasks given equal weight in the aggregate. On that average, LM-X had 74.1% mean success versus 55.4% for GR00T N1.7, an 18.7-point absolute difference. The representative-task comparison was not uniformly favorable.

The physical-robot comparison covered seven tasks on four embodiments, using 20 trials per task. The methods used identical post-training data and success criteria. LM-X recorded 68.6% mean success versus 50.7% for GR00T N1.7, a 17.9-point difference, although the comparison was not uniformly favorable across tasks.

Readable patterns, but no tested alarm

Held-out RTG traces generally increased as episodes moved toward completion, changed near semantic events and decreased rapidly at visible anomalies. That is temporal correspondence with task progress, not calibrated detection accuracy.

Variance traces showed repeated spikes in failed episodes, while successful episodes tended to remain lower and smoother. Local increases appeared during oscillation and fell after the robot committed to a movement. These observations describe how the signal changed during episodes, not a calibrated threshold for detecting failure.

What the comparisons leave open

The reported differences also carry statistical uncertainty. The simulation study used a single training run for each configuration, while the real-world evaluation was limited to 20 trials per task and reported no confidence intervals.

The comparisons do not isolate the contributions of the architecture, the pretraining data or any individual predictive head. The RTG and variance analyses likewise do not establish calibrated detection performance; labeled failure windows, threshold selection and held-out calibration are still needed.

Taken together, the study reports higher mean success for LM-X in the tested gate, simulation aggregate and real-world aggregate, along with recognizable episode-level patterns in RTG and variance. It does not show that the predictive heads caused the gains or that the online signals can reliably detect failure. The supplied document is an arXiv preprint, version 2, dated 27 Aug 2026.

Paper data and sources

Original title: LM-X: Explainable Action Modeling with Progress, Event, and Uncertainty Prediction for Generalist Robot Manipulation
Authors: Jin Lou, Jingxuan Zhu, Andong Chen et al.
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-26
DOI: Not available
Original paper · Full text

Versions and corrections

  1. Published automatically after legal-source, freshness, evidence, and independent-verification gates passed.