A robot system that reads how a compliant gripper deforms during contact reported stronger results than visual and visuo-tactile comparison systems across three tasks: grasping objects at different scales, unscrewing caps and writing calligraphy. In the standard tests, it recorded 100% successful grasps with no fall-offs, 90% successful cap-unscrewing runs with no dislocations, and 100% holding success with 90% writing success.
The work is an arXiv version-1 preprint dated 26 Aug 2026. Its evaluation used a UR3 arm, a UMI-based passive compliant gripper and RealSense L515 and D405 cameras, with expert demonstrations and robot trials rather than human or animal participants.
How the robot reads contact
VISTA-Policy uses a three-dimensional Visual Deformation Field, or VDF, as visuo-physical feedback. The field captures the gripper's deformation and is processed by a Physics-Aware Encoding Engine, an Energy Aggregation Denoising Mechanism and a Deformation-Augmented Policy Network. The policy also uses incremental gripper actions, updating movement in relative steps.
The comparison included DP3, which uses global point clouds; DP3-Wrist, which uses a wrist-view point cloud; and TDF-DM, which uses a commercial Daimon visuo-tactile sensor. Other tests changed the action representation, removed denoising from deformation processing or replaced the encoder with an MLP. Because the full VISTA package combines several of these choices, the study does not isolate VDF alone as the explanation for every difference.
The strongest numbers came in standard trials
On grasping, VISTA's success rate was 100% and its fall-off rate was 0%, compared with 40%/70% for DP3, 50%/50% for DP3-Wrist and 60%/60% for TDF-DM. For cap unscrewing, VISTA recorded 90% success and 0% dislocation, while the corresponding pairs were 50%/20%, 40%/30% and 50%/80%.
Calligraphy produced 100% holding success and 90% writing success for VISTA. The comparison figures were 50%/40% for DP3, 90%/40% for DP3-Wrist and 80%/20% for TDF-DM; an MLP-based VISTA variant reported 100%/60%. Holding success and writing success were tracked separately, so the two outcomes were not interchangeable.
The pattern continued in harder tests
In single-object cross-scale tests, VISTA reported 100% success and 0% fall-offs in both listed conditions. The absolute-action VISTA variant also began at 100%/0% but fell to 40% success with 60% fall-offs in the second condition, while DP3-Wrist recorded 60%/40% and 50%/50% across the two conditions.
With multi-scale training, VISTA reported 100% success and 0% fall-offs in both seen and unseen multi-object conditions. DP3 reached 40%/60% in the seen condition and 40%/80% in the unseen one; DP3-Wrist reached 60%/40% and 40%/60%. The VISTA-Cat ablation, which concatenated deformation without denoising, reported 80%/80% and then 20%/80%.
Cap unscrewing also remained strong at boundary scales: VISTA's four reported success/dislocation pairs were 100%/0%, 80%/20%, 80%/40% and 80%/0%. DP3-Wrist's corresponding pairs were 20%/40%, 20%/60%, 60%/20% and 40%/20%. In multi-object training, VISTA was 100%/0%, compared with 50%/20% for DP3, 40%/30% for DP3-Wrist, 60%/40% for VISTA-Cat and 50%/80% for TDF-DM.
In the grasping sample-efficiency analysis, DP3-Wrist trained with 40 demonstrations remained below VISTA trained with 20. The baseline also over-gripped medium-to-large objects, according to the reported analysis.
Recovery and unfamiliar heights
During a disturbance test, an object was manually knocked back while being lifted. VISTA completed a secondary recovery grasp in four of five trials, or 80%, while DP3-Wrist failed completely. Qualitative tests also reported adaptive-compliance manipulation of tofu and cards.
At out-of-distribution writing heights, VISTA maintained paper contact and short-stroke writing at a lower bound of 15 cm, and adapted when the surface height changed online. DP3-Wrist had lower-height contact failures and lost control stability after height transitions.
What the signal analysis showed
Feature analysis found high similarity among VDF features when local contact states were similar, even across different object geometries and scales. It also showed clearer separation between non-contact and progressive-contact phases than DP3 point-cloud features. The authors interpret those patterns as evidence of a more structured physical representation, but the analysis is not an independent measurement of force or compliance.
The paper reports a $77 compliant gripper that endured more than 1,000 high-intensity, contact-rich trials without structural degradation.
What the results do not establish
The study's default plan used 20 expert demonstrations, 10 standard trials per configuration and five robustness or OOD trials per configuration. The supplied analysis reports descriptive percentages without confidence intervals, hypothesis tests or p-values.
Those limits matter for interpretation. The evidence comes from one robot, one compliant-gripper design, one camera arrangement and the listed tasks and conditions, so results on other robots, grippers, cameras, environments or broader contact-rich tasks remain untested. The current deformation representation is mainly contact-state feedback, not demonstrated force or compliance estimation.
Paper data and sources
Original title: VISTA: Visually Inferred Spatial ConTact Attention for Contact-Rich Manipulation
Authors: Jiayi Chen, Wenlong Dong, Yan Huang et al.
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-26
DOI: Not available
Original paper · Full text