One model, three instruction modes
A preprint reports that a single pretrained robot model can support three ways of giving instructions: language alone (LI), language plus a deictic gesture (VLI), and a gesture alone (VI). The design turns each mode into a shared representation made of a text prompt and deictic masks. A deictic mask is the part of the input that marks the target indicated by the gesture, allowing the same model to receive words, a target cue, or both.
The paper tests how gesture information reaches the model. VP-BBox and VP-Fade are RGB visual-prompting methods, while MP-Early and MP-Late pass the mask through a separate channel. Its main training plan has two stages: first adapting the model with LI data, then jointly training on LI, VLI and VI. The study also compares other stage choices and whether LI data stays in the second stage.
Strong scores in simulation
In simulation, each task supplied 50 expert demonstrations and was evaluated in 25 trials, using task success rate. On the two-stage, in-distribution comparison, the four prompting methods all reached mean success rates between 94.1% and 95.6%. VP-BBox was highest at 95.6%.
The sharper differences appeared in zero-shot tests, where the paper measured how much success changed when valid deictic-mask input was available versus when the mask-derived input was disabled. For Object-ZS, VP-BBox's difference was +29.2 percentage points for VLI and +27.8 for VI. MP-Late recorded +29.2 and +25.2. For Spatial-ZS, VP-BBox recorded +16.0 for VLI and +11.1 for VI.
With cumulative training steps matched, the paper associated two-stage training mainly with deictic-mask use in unseen layouts, rather than with a clear improvement in in-distribution task success. The pattern depended on the prompting method, so the simulation does not identify a universal winner.
A separate two-stage setup omitted LI data from stage two. There, LI success was 63.6% to 72.1%, while VLI and VI remained between 95.4% and 97.0% in-distribution. The reported comparisons came without confidence intervals or inferential tests, so they describe observed success rates rather than a quantified level of statistical certainty.
Gesture modes held up better in real-world tests
The physical-robot evaluation used one policy trained to support all three modes, with 100 to 150 episodes per task. The demonstration set contained 360 episodes and about 97,000 recorded timesteps. In the reported training-distribution comparison, VLI success ranged from 50.0% to 77.5%, VI from 45.8% to 75.0%, and LI from 29.2% to 62.5%.
The difference widened when the wording was unfamiliar. At unseen instruction-expression levels TL2 and TL3, PickBlock produced LI success rates of 12.5% and 18.8%, compared with 62.5% and 68.8% for VLI and 62.5% and 93.8% for VI. In PutBlock, the corresponding figures were 26.7% and 20.0% for LI, 86.7% and 80.0% for VLI, and 86.7% and 75.0% for VI.
Other unseen conditions showed the same direction. In OrganizeToy, under the unseen table-surface condition, jointly trained LI reached 60.0%, while VLI reached 87.5% and VI 90.0%. For novel object instances, VLI and VI each reached 95.8%, versus 62.5% for LI. For novel object categories, VLI and VI each reached 100%, compared with 16.7% for LI.
What the study does not settle
Taken together, the results support the authors' view that VLI and VI can complement LI when language-only target specification is unreliable. They do not establish that gesture-based instructions are always better, that one prompting method dominates across all conditions, or that the training strategy caused the observed differences. The study reports descriptive success-rate comparisons, with no confidence intervals or inferential tests.
That caution matters because the work is still a version 1 arXiv preprint dated 28 August 2026. It was supported by JST Moonshot R&D Program Japan Grant JPMJMS2011 and JST BOOST, Japan Grant JP-MJBS2402.
Paper data and sources
Original title: DeicticVLA: Unifying Instruction Modes Based on Language and Deictic Gestures in a Single VLA
Authors: Kango Yanagida, Tatsuya Aoki, Yuichiro Yoshikawa, Takato Horii
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-28
DOI: Not available
Original paper · Full text