A preprint reports that a robot policy performed better than two comparison systems when asked to carry out manipulation tasks it had not practiced during training. In seven unseen RoboTwin 2.0 simulation tasks, Zero-WAM recorded an average success rate of 46.95%, compared with 17.45% for LingBot-VA and 10.98% for WAN-Action. That amounted to a lead of 29.50 percentage points over LingBot-VA and 35.97 points over WAN-Action.
The simulation comparison was split at the task level: 43 tasks were used for post-training and seven were held out as unseen evaluation tasks. Each unseen task was run through 100 closed-loop rollouts under each of three random seeds, so the reported averages summarize repeated simulated runs rather than a single attempt.
A video as the task specification
Zero-WAM is a causal video-action policy that supports both language instructions and human video demonstrations as task specifications. It predicts future robot video and executable actions. The study asks whether that deployment-time information is enough to execute a manipulation task not practiced during training.
Those human-video prompts came from a pipeline called HumanGen. The generated videos vary the background, camera viewpoint, environmental style, object instance and object placement while preserving the task semantics of the robot trajectory. HumanGen contains 74.2K human-robot in-context-learning pairs spanning 8.6K tasks.
Training for a larger task mix
Zero-WAM was trained with a task-diverse video-action set that sampled more than 6,000 tasks and approximately 400K robot trajectories per training epoch. The training also included IFP, or in-context future chunk prediction, an auxiliary objective used only during training. IFP asks the model to predict multiple strided future robot-video chunks from its current representation, with the stated aim of retaining longer-term task evolution from the human video.
The advantage was spread across the simulation set rather than driven by a single task. Zero-WAM outperformed both baselines on all seven unseen tasks, reached 84.87% on place empty cup, and was the only method in the main comparison to register any success on stack blocks three.
The real-robot test was narrower
The real-world evaluation used a bimanual Franka and a 252-pair in-context dataset. It included 120 pairs from 30 placement combinations, 96 pairs from 16 sequential-manipulation combinations and 36 insertion pairs. For unseen task configurations, the paper reports that the evaluation used no corresponding robot data and did not update the model parameters.
Over 30 real-robot trials, Zero-WAM's success rate was 53.3% for placement, 33.3% for sequential manipulation and 16.7% for insertion. LingBot-VA's corresponding rates were 43.3%, 10.0% and 0.0%. The percentages point in the same direction as the simulation comparison, but they describe three listed task families rather than a broad test of real-world manipulation.
What the ablations add
The ablation results were consistent with a role for in-context human-video conditioning. In that comparison, the reported average success rate was 36.36%, versus 17.45% for LingBot-VA and 10.98% for WAN-Action. That is a higher score for the video-conditioned configuration, but it does not by itself establish that video alone caused the difference.
IFP was examined separately. The seven-task average was 46.95% with IFP and 28.55% without it; on stack blocks three, the corresponding figures were 9.00% and 0.00%. These are within-study comparisons of the model variants, not a guarantee that the same gap would appear on every new task.
Task-balanced pretraining was also tested with a text-only Zero-WAM variant. That version reached a 39.44% average across the seven unseen tasks, 21.99 percentage points above LingBot-VA. The result suggests that the reported advantage was not confined to human-video prompts.
A result with clear boundaries
How much confidence to place in the numbers depends on the test size and design. The simulation averages were summarized across three random seeds, while the real-world percentages came from 30 trials; the supplied real-world results include no variability estimate.
The evidence stays within seven simulated evaluation tasks and three listed real-world task families on a bimanual Franka. It does not establish that the policy will handle arbitrary open-ended tasks, or dynamic and mobile environments beyond this setting.
The document is an arXiv version-one preprint dated 26 August 2026.
Paper data and sources
Original title: Zero-WAM: In-Context World-Action Modeling from Human Videos for Open-Ended Task Generalization
Authors: Jiaming Zhou, Qihang Zhang, Gangwei Xu et al.
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-26
DOI: Not available
Original paper · Full text