Preprint

MA-VLA reports higher scores on specified multi-arm collaboration patterns

Preprint: MA-VLA reported higher benchmark scores than Pi0 and nonzero results on specified collaboration patterns, with tests limited to selected simulations and one dual-arm platform.

A robotics preprint reports that MA-VLA scored higher than Pi0 on the listed in-domain benchmarks and recorded nonzero success when known atomic actions were recombined into specified multi-arm collaboration patterns absent from training. The tests covered two simulation benchmarks and one real dual-arm robot platform, so the results describe selected evaluations rather than open-ended coordination.

The system divides a shared instruction

MA-VLA combines a vision-language model planner with a vision-language-action executor. The planner turns an instruction into temporally ordered atomic subgoals for each arm, while the executor grounds those subgoals in robot actions.

The study asks whether the system can recombine known atomic actions into collaboration patterns absent from training. The experiments compare Clean training, without Arm Shuffle or View Dropout, with Regularized training, which enables both augmentations.

Higher averages on the listed benchmarks

On RoboFactory, MA-VLA reported averages of 83.5% and 83.3% in the listed task blocks, compared with 80.3% and 76.5% for Pi0. On RoboTwin 2.0 (Hard), the reported values were 49.0 for MA-VLA and 41.1 for Pi0.

For simulation training, the data included 150 expert demonstrations for each task. Each simulation configuration was then evaluated over 100 rollouts using mean task success rate.

The tougher comparison used specified combinations

In simulation tasks involving collaboration patterns absent from training, all listed baselines reported 0.0 average success, while MA-VLA reported 13.0. Its reported task values included 28.0 for GBR, 9.0 for three listed tasks and 10.0 for Stack Two Bowls.

On the real-world test's four dual-arm tasks on the SO101 platform, the out-of-domain tests used specified collaboration patterns outside training. MA-VLA reported 10 successful episodes out of 20 on one, 8 out of 20 on another, and 2 out of 20 on each of the remaining two; Pi0 reported 0 out of 20 on every out-of-domain task. MA-VLA also reported higher in-domain counts on all four tasks.

Each real-world task used 50 teleoperated demonstrations, followed by 15,000 training steps with a batch size of 32 and 20 evaluation episodes.

The reported gains varied across configurations

An ablation reported 0.0% out-of-domain and 48.0% in-domain success without added components. The atom-only configuration reported 0.0% and 58.0%; atom plus shuffle reported 7.3% and 52.0%; and the full-component configuration reported 15.3% out-of-domain and 53.0% in-domain success.

A separate-arm Pi0 setup used three VLA models and reported 0.0% out-of-domain, 61.0% in-domain and 30.5% average performance. MA-VLA used one model and reported 15.3% out-of-domain, 53.0% in-domain and 34.2% average performance. The reported shuffle-rate analysis found higher out-of-domain performance at larger shuffle rates, while in-domain accuracy showed a mild drop and remained stable.

What the results do and do not cover

The out-of-domain evaluations used predefined recombinations of known atomic actions and collaboration patterns. They therefore do not establish generalization to arbitrary collaboration structures or transfer beyond the tested simulation benchmarks, tasks and SO101 platform.

The results are aggregate success values or counts, and the analysis includes no confidence intervals, standard errors or hypothesis tests. The planner uses GPT-4.1, a predefined atomic-prompt vocabulary and prompting choices, while baseline interfaces were adapted for multi-arm inputs.

The paper is an arXiv version 1 preprint dated 26 August 2026. It reports support from the National Natural Science Foundation of China, the Liao Ning Science and Technology Plan and the Dalian City Science and Technology Innovation Fund, and states that code, models and data are available at a listed GitHub repository.

Paper data and sources

Original title: MA-VLA: Multi-Arm Vision-Language-Action Model for Collaboration and Compositional Generalization
Authors: Zaibin Zhang, Junlan Xiao, Zhongbo Zhang et al.
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-26
DOI: Not available
Original paper · Full text

Versions and corrections

  1. Published automatically after legal-source, freshness, evidence, and independent-verification gates passed.