A preprint analysis suggests that physical-AI leaderboard positions can depend heavily on how overlapping tests are counted. When researchers collapsed two pairs of near-substitute benchmarks into pair averages, 22 of 51 models shifted by at least three places in an arithmetic-average ranking.
The audit used a matrix of 51 models selected from a registry of 152. Each selected model had at least 8 of 12 benchmark scores reported. Scores came from model cards and papers, supplemented by 159 data points from targeted author runs using official evaluation code and prompts when available; the authors did not conduct a systematic reproduction study.
Where overlap changes the story
The clearest overlap appeared in two benchmark pairs. Spearman correlation, a measure of how similarly two tests rank models, reached 0.876 for EmbSpatial and CV-Bench across 49 models, with a reported 95% confidence interval of 0.78 to 0.93. Where2Place and RefSpatial-Bench correlated at 0.860 across 50 models, with a reported interval of 0.73 to 0.93. Both pairs cleared the 0.8 threshold the analysis used to flag close substitutes.
A separate predictability check modeled each benchmark from the other 11, using ridge regression and leave-one-model-out cross-validation. Where2Place, RefSpatial-Bench and ERQA were the most predictable from the rest of the suite, with R2 values of 0.727, 0.721 and 0.706. The authors note that low predictability can point either to genuinely distinct information or to measurement noise.
Those choices have practical consequences for aggregate leaderboards. After the two substitute pairs were collapsed, 22 of the 51 models moved by three or more positions. The paper does not present the resulting de-duplicated order as uniquely correct, because deciding which benchmarks count as substitutes is itself a judgment.
A common signal across the suite
The analysis also found a broad common pattern across the 12 physical-AI benchmarks. Its first principal component, a statistical summary of the strongest shared variation, explained 55.2% of the total variance. Every benchmark loaded positively on that component, with loadings ranging from 0.48 to 0.88. That is a description of how the scores move together, not evidence that one capability causes another.
This physical-benchmark axis closely tracked the first component extracted from the general anchors. The Spearman correlation was 0.952; correlations with MMStar and Video-MME were 0.950 and 0.942. The pattern is consistent with a substantial general-capability component in the shared scores, although the observational design cannot establish a causal explanation for it.
The internal agreement weakened once the external general-capability axis was removed statistically. Mean pairwise absolute correlation fell from 0.487 to 0.250, while the number of benchmark pairs above 0.5 dropped from 34 of 66 to 6. The result depends on how the external anchors were constructed and on which models had overlapping scores.
A shorter suite, with a clear caveat
The researchers then searched for a smaller set that could still separate models while adding information not already recoverable from the chosen tests. Their utility measure combined Gini score dispersion, which rewards spread in model scores, with the share of candidate variance that could not be reproduced from the selected set.
A greedy selection path chose RefSpatial-Bench, MindCube, VSI-Bench and BLINK as its first four benchmarks. Together, those four retained 78.5% of the utility measured for all 12 benchmarks.
Sensitivity checks that forced different benchmarks to start the selection still kept MindCube and VSI-Bench in the first four in every run, while RefSpatial-Bench was never excluded. Reported four-benchmark utility ranged from 69.2 to 78.5. The authors describe the set as one defensible choice, not a guaranteed global optimum.
Using the compact set as the basis for a Bradley-Terry leaderboard, HY-Embodied-0.5 MoE-407B-A32B ranked first with an Elo score of 2251, followed by Qwen3.5-397B-A17B with 2032. The displayed table covered the top 10 of the 51 models. The ranking is an observational fit to benchmark outcomes, and models had varying coverage of the selected tests.
What the analysis cannot answer
The study does not show that general capability causes performance on physical-AI benchmarks, or that the remaining benchmark signal transfers to manipulation or navigation. None of the 12 benchmarks measures downstream success on a physical system, so a compact leaderboard should not be read as a direct test of robot task performance.
The findings are tied to the selected model-by-benchmark matrix and the particular utility definition. The scores came from mixed reported and author-run sources, and no systematic reproduction study was performed, so the analysis does not settle how much of the apparent uniqueness reflects measurement noise.
The supplied document identifies itself as arXiv version 1, dated 26 August 2026.
Paper data and sources
Original title: A Statistical Audit of Physical AI Benchmark Redundancy
Authors: Zaruhi Navasardyan, Hrant Davtyan
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-26
DOI: Not available
Original paper · Full text