The preprint reports that DECOWAM was more accurate than a 50k-step FastWAM starting point at predicting future video and whole-body robot actions in a fixed replay test. Frame mean-squared error was 15.03% lower, action mean-squared error was 21.71% lower, and PSNR was 0.222 decibels higher.
The model is designed to distinguish camera ego-motion from movement of the robot’s base and arm while jointly predicting future observations and actions. The reported physical tests used a quadruped-arm platform.
A model built to separate movement from manipulation
DECOWAM freezes an adapted FastWAM backbone and adds residual adapters, a bottleneck that limits access to privileged future information, separate base and arm representations, and conditioning on base velocity for video prediction.
The quality-filtered ARMDOG corpus contains 1,487 episodes, 343,550 RGB frames and 321.3 minutes of data recorded at 15 frames per second. Stage 2 used 214 episodes from 26 tasks, while the replay evaluation used a fixed slice of 23 episodes covering eight tasks and 4,323 frames.
Strong scores, with one model still ahead
Among the WAM reference systems, DECOWAM ranked first on every reported video and action metric. Its frame and action errors were also reported as lower than those of FastWAM, Motus and X-WAM.
X-VLA had the lowest error on all three reported action measures. DECOWAM ranked second on action mean-squared error and action L2 error; those errors were 69.9% and 21.1% lower than π0.5. Its action mean absolute error was within 11.1% of π0.5 and 84.0% lower than GR00T.
The component tests showed the same reported pattern: every module-removal version performed worse than the full model. Compared with an adapter-only control, the full system’s frame error, action error and action mean absolute error were 4.6%, 15.7% and 16.1% lower, respectively.
Better coordination did not mean reliable completion
In closed-loop trials, each method was tested 79 times under common observation, language-input, low-level-control and safety settings. DECOWAM completed 46 tasks, with a 58.2% success rate and a mean completion time of 49 seconds.
DECOWAM recorded the highest reported approach and transport rates, and 96.4% of successful grasps carried over into transport. It also tied the highest docking rate and led the reported coordination, displacement-robustness and recovery measures.
Compared with FastWAM, the reported coordination and displacement-robustness rates were 10.1 and 17.7 points higher. Compared with X-WAM, the differences were 16.5 and 25.3 points.
A smaller update footprint, but narrow evidence
During Stage 2, DECOWAM updated about 25.95 million parameters, compared with 6,020.75 million for FastWAM—an approximately 232-fold smaller update footprint. Evaluator latency was reported as 11.4% higher.
The figures need a narrow reading. The main replay comparison used a fixed 23-episode slice, and the physical evaluation used 79 trials per method. The supplied analysis reports no inferential tests, confidence intervals, p-values or cross-seed uncertainty estimates.
These are reported differences in the tested ARMDOG setting, not proof that DECOWAM caused the gains or that they will hold on other platforms and task distributions. The parameter result also concerns Stage 2 updated parameters, not the size of the entire deployed network.
The paper describes a planned release of raw and cleaned HDF5 files, the world-action conversion, immutable split manifests and a datasheet. It does not state that those materials are already available.
Paper data and sources
Original title: DECOWAM: Decoupled Whole-Body World-Action Model for Legged Mobile Manipulation
Authors: Siyuan Ma, Boshi Zhang, Yutian Zhang et al.
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-20
DOI: Not available
Original paper · Full text