A robot-grasping model led the compared methods in simulated tests while using five network evaluations, according to an arXiv preprint posted on Aug. 26, 2026. At that setting, GraspMF had the highest reported success rates for both in-domain objects and objects from disjoint categories, with a reported inference latency of 15.5 milliseconds.
GraspMF applies MeanFlow on SO(3) × R^3, representing each grasp through rotation and translation. It combines an algebraic semigroup-consistency condition with a Riemannian Conditional Flow Matching anchor tied to the data distribution.
The benchmark result
The main benchmark used 416 ACRONYM object instances. For in-domain evaluation, 90% of the instances in each category were used for training and 10% for evaluation; out-of-domain evaluation used disjoint categories. In simulation, 100 grasps were sampled per object, and a grasp counted as successful if it retained the object after a five-second shake.
At five evaluations, GraspMF reported 87.4% success for in-domain objects and 71.73% for out-of-domain objects. Its Earth Mover's Distance (EMD), a measure comparing sampled grasps with reference grasps, was 0.3702 in-domain and 0.4191 out-of-domain; it had the best reported out-of-domain EMD and a competitive in-domain value. The closest baselines were 1.9 to 2.7 percentage points lower on success rate and 0.02 to 0.05 higher on EMD.
Speed at a lower sampling budget
At the reported settings, GraspMF used five network function evaluations at T = 5, compared with 140 for SE3Dif, 140 for VSIGD and 80 for EGF. Its reported latency was 15.5 ± 0.4 milliseconds for a batch of 100 grasps on one NVIDIA RTX 5080 GPU; the abstract reports a speed-up of up to 39 times.
At one evaluation, GraspMF reported 81.11% in-domain and 66.34% out-of-domain success, with EMD values of 0.3698 and 0.4218, one network function evaluation and latency of 6.3 ± 0.2 milliseconds. Across the tested budgets of 1, 2, 5, 10, 20, 40, 70 and 100, out-of-domain success averaged 70.15%, with a standard deviation of 1.63 percentage points; baseline performance degraded sharply at low budgets.
A harder view and a physical demonstration
To test partial observations, the model was given a single view at a time, with results averaged over three random views per object. At five evaluations, GraspMF recorded 86.08% in-domain success and 68.46% out-of-domain success, with EMD values of 0.4030 and 0.5923. BRIDGER recorded 88.84% in-domain and 65.38% out-of-domain success, giving it the higher in-domain result and GraspMF the higher out-of-domain result.
In a separate physical demonstration, the system was tested on a black mug, a red mug and a gray bowl, with 10 trials for each object at each setting. At five evaluations, it recorded 9 of 10 successes for each mug and 10 of 10 for the bowl. At 10 evaluations, it recorded 9 of 10 for the black mug and 10 of 10 for the red mug and bowl.
What the evidence leaves open
The physical demonstration does not establish how the model would perform beyond those three objects. More broadly, the evidence is limited to the reported object categories, simulated environments, synthetic partial views and the stated physical setup; it does not validate end-to-end closed-loop replanning.
The partial-view experiment was synthetic: observations were created by raycasting, sensor noise was not quantitatively controlled or characterized, and the EMD reference set included only grasps contacting observed surfaces. Those choices complicate interpretation of the partial-view EMD. Direct comparison with the baselines is further limited because the systems used different architectures, network-function evaluation schemes and reported sampling budgets.
Latency was measured on one GPU with one batch configuration, so the reported latency comparison may not carry over to other hardware, batch sizes or complete closed-loop execution. The preprint reports no confidence intervals, formal hypothesis tests or sample-size justification, and the statistical meaning of the reported plus-or-minus figures is not specified.
In component-removal tests, the variant using a decomposed MeanFlow objective instead of the semigroup-consistency loss was associated with the largest reported degradation: 25.15 percentage points in-domain and 19.44 out-of-domain. At five evaluations, the ablation without signed-distance regression was associated with in-domain and out-of-domain success decreases of 3.8% and 5.1%, while the ablation without the schedule was associated with decreases of 2.4% and 5.8%.
The Gram–Schmidt rotation-projection variant reported 13.5 milliseconds of latency versus 15.5 for GraspMF and 87.51% versus 87.40% in-domain success. Its out-of-domain success was 4.7% lower than GraspMF's, and its EMD was higher in both domains.
The abstract reports sampling in five or fewer network evaluations, millisecond-scale latency, a speed-up of up to 39 times and direct transfer to physical grasping without additional training or domain adaptation. In the reported work, the evidence for those claims comes from the stated benchmarks, one GPU and a three-object demonstration; whether the results persist across broader categories, robots, sensors, hardware and controlled observation noise remains open.
Paper data and sources
Original title: Fast Generative Grasping via Lie Group-Constrained MeanFlow
Authors: S. Talha Bukhari, Yi Wei, Ruiqi Ni et al.
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-26
DOI: Not available
Original paper · Full text