A robot system that tells an action model where to grasp or place an object using visual geometry rather than generated coordinate text recorded higher completion rates in Bridge/WidowX and physical-robot tests. Pointing-VLA paired with CuRobo completed 72.9% of Bridge/WidowX tasks on average, compared with 52.1% for Embodied-R1 paired with CuRobo. In physical AgileX PiPER trials across three visual contexts, π0.5 + P-VLA completed 121 of 150 attempts, or 80.7%, against 79 of 150, or 52.7%, for π0.5. The reported difference was 28.0 percentage points.
A spatial interface for robot instructions
Pointing-VLA is built as an interface between a vision-language-action model and the machinery that executes a robot move. It reads spatial intent from multimodal hidden states — internal representations formed while the model processes its inputs — and turns that information into three outputs: a point, an OFG/contact heatmap for contact or affordance, and a visual trajectory. It binds each structured execution slot to the geometry that slot needs. In the tested contract, source-conditioned OFG supplies the pick location and Pointing supplies the place location.
Training follows the same logic: point, region and visual-trace examples supervise their corresponding learned heads. The evaluation scored Pointing and OFG with point-in-mask or point-in-box success, visual trajectories with normalized RMSE, ADE and FDE, and robot studies by full task completion. Within each experiment, compared methods used the same samples or episode identities. The Bridge/WidowX deployment covered four manipulation structures with 24 episodes per task.
The readouts did not have a single winner
On native grounding tests, Pointing matched Embodied-R1 at 64.3% sparse referring-point accuracy. OFG showed a different pattern on the full Part-Affordance-2K task: it reached 57.3%, versus 40.9% for Embodied-R1, a reported difference of 16.4 percentage points.
That crossover appeared across shared samples from other datasets. Pointing was strongest on RefCOCO and RoboAff, while OFG was strongest on AGD20K. The preprint describes that pattern as supporting expert-specific output geometries.
On 300 VABench-V examples, the visual-trace readout recorded 0.1042 normalized RMSE, 0.1368 ADE and 0.1493 FDE in normalized image coordinates. In a shared runtime protocol, the geometric readouts were reported as 6.68 to 6.90 times faster than autoregressive text decoding, the step-by-step generation of a coordinate string.
A fixed contract carried into manipulation tests
One evaluation focused on source selection: identifying which object should be picked. In the frozen multicolor Stack cases, source-conditioned OFG selected the instructed source in all 48 online cases. With identical Pointing PLACE targets, it completed 43 of 48 episodes, compared with 40 of 48 for Attention PICK.
Across the phase-aligned Bridge/WidowX deployment, the fixed OFG-PICK/Pointing-PLACE contract completed 70 of 96 episodes, or 72.9%, under collision-enabled CuRobo. In that deployment, it completed all Stack episodes and 22 of 24 Eggplant episodes. By task, Pointing-VLA + CuRobo recorded 50.0% success for Spoon, 50.0% for Carrot, 100.0% for Stack and 91.7% for Eggplant.
The physical trials showed a similar gap
On the AgileX PiPER, the physical pick-and-place comparison used three visual contexts with 50 trials each. With P-VLA, π0.5 completed 36 of 50 trials (72.0%) in the no-distractor context, 42 of 50 (84.0%) with a yellow cylinder and 43 of 50 (86.0%) with a red cylinder. Without P-VLA, π0.5 completed 20 of 50 (40.0%), 26 of 50 (52.0%) and 33 of 50 (66.0%) in those contexts. Across all 150 trials, the totals were 121 successes (80.7%) with P-VLA and 79 (52.7%) for π0.5.
The failure counts followed the same pattern: grasp failures were recorded 47 times for π0.5 and 16 times for π0.5 + P-VLA, while tray failures were recorded 13 and four times, respectively.
Transfer kept the main model unchanged
A transfer check kept the NORA-1.5 backbone frozen while moving an OFG/contact readout into the shared wrapper. Horizontal-pose success was preserved. For laid-vertical poses, success was 95.0% with the transferred readout versus 89.0% for the NORA base, and recorded controller time was reported as more than 20 times shorter.
The document is an arXiv preprint, version 1, dated 24 Aug 2026. Its evaluation was built around four questions: whether typed heads preserve task-appropriate geometry; whether source-conditioned OFG resolves pick-source selection within the fixed OFG-PICK/Pointing-PLACE contract; whether the readout transfers while reducing inference cost; and whether performance holds in simulation and physical-robot deployment.
Paper data and sources
Original title: Pointing-VLA: Typed Spatial Grounding Interfaces for Vision-Language-Action Manipulation
Authors: Xiwen Chen, Zelin Li, Zhiruo Zhou et al.
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-24
DOI: Not available
Original paper · Full text