Preprint

A Robot Grasped Objects With Partial 3D Reconstructions in Lab Tests

Preprint reports 65.6% success across 192 grasp attempts, while a low-resolution method used fewer views than a standard approach.

A robot in a laboratory packing test succeeded in 126 of 192 grasp attempts, or 65.6%, while reconstructing only part of an object's surface. It failed 66 times, or 34.4%. Mean reconstruction at grasp time ranged from 39.79% to 59.83% across the tested object and policy-resolution conditions.

The work is a preprint identified as arXiv:2608.25874v1 and dated 26 Aug 2026. Its front matter says it has been accepted to IFAC for publication under a Creative Commons CC-BY-NC-ND licence.

Fewer views, comparable reconstruction

Study A focused on how the robot chose where to look. Each of four objects was reconstructed five times with all the methods. LR-NBV was compared with standard NBV using the same Exploration gain but omitting Density gain; both used the same stopping rule based on marginal gain. HEU served as a second reference and sampled 13 viewpoints while assuming the object's pose was known.

To judge the reconstructions, the researchers used bounding-box dimension intersection-over-union, or IoU, a measure of agreement between an estimated box and the reference. LR-NBV had statistically significant IoU gains for the horizontal box and the gripper. For the vertical box and the flange, significant differences appeared in at least one bounding-box dimension. The reported cutoff was p<0.05, and the comparisons used the Mann–Whitney U test across methods, metrics and objects.

The efficiency result was more pronounced. LR-NBV consistently required fewer poses than NBV, often fewer than half as many, while maintaining comparable reconstruction percentages. NBV also produced more useless poses, and the reported difference in pose counts met the p<0.05 threshold.

Results shifted with the setup

Study B tested the full pipeline in a 2 × 2 design, crossing the R60 and R30 input resolutions with the XPLT and XPLR policies across two object types. It ran 16 trials per setting, eight per object, for 64 trials and 192 grasp attempts in all, with no re-attempts.

The hardware was an ABB GoFa 12 arm paired with a MaixSense A010 time-of-flight depth sensor. The inputs were downsampled to R60 and R30, and depth-only GraspNet generated six-dimensional grasp hypotheses.

Success rates varied by object and setting. For Obj A, rates in R60–XPLR, R60–XPLT, R30–XPLR and R30–XPLT were 62.5%, 70.8%, 54.2% and 54.2%, respectively. For Obj B, the same four settings produced 66.7%, 58.3%, 83.3% and 75.0%.

At grasp time, mean reconstruction for Obj A was 42.55%, 56.29%, 39.79% and 44.36% in that order. For Obj B, it was 54.77%, 59.83%, 50.45% and 47.32%. The figures show that the tested policy could act without an exhaustive surface reconstruction.

Failures labeled STAB accounted for 33.3% of all failures. The average STAB share was 37.5% under XPLR and 29.4% under XPLT.

The acquisition curves also differed by policy: XPLR followed a clearer logarithmic pattern, while XPLT was closer to linear.

What the experiment does not settle

The evidence is narrow. Study A used four objects, each reconstructed five times, while Study B tested two object types in 64 trials, 16 per setting and eight per object, with no re-attempts. The hardware and resolutions were also fixed to the tested setup.

That means the results describe these combinations rather than a broad range of packing conditions. The condition-level success rates ran from 54.2% to 83.3%, and the mean surface reconstructed at grasp time stayed between 39.79% and 59.83%.

The statistical picture is limited too: the paper reports the p<0.05 threshold but not exact p-values or effect sizes for Study A, and gives no confidence intervals for Study B's overall or condition-specific success rates. Exact pose counts are not reported either.

The authors acknowledge Camozzi Research Center for providing robotic equipment and support.

For now, the claim is a limited one: in the tested setups, a robot recorded successful grasps from partial reconstruction, and LR-NBV used fewer views than NBV. The experiments do not establish performance beyond the tested robots, sensors, objects, policies and resolutions.

Paper data and sources

Original title: Low-Resolution Perception for Robotic Packing
Authors: Giuseppe Fabio Preziosa, Federico Vignoni, Chiara Castellano et al.
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-26
DOI: Not available
Original paper · Full text

Versions and corrections

  1. Published automatically after legal-source, freshness, evidence, and independent-verification gates passed.