A smart-home simulator pass did not match every physical check in an arXiv preprint audit. In a 240-trial light-calibration block, a false clearance meant that a source-cleared case failed the matched physical check; the counts were 240/240 for completion, 20/240 for reported read, 42/240 for observable effect and 0/240 for settled control.
The results changed with what was observed and when it was read. Across two devices and a grid extending to 0.05 seconds, settled effect was uniformly valid; observable failures were off-only through 0.25 seconds, reported state had a sub-50-millisecond boundary, and completion was uniformly invalid at its declared read.
The audit compared matched traces
SimVerity reran declared scenarios on the target deployment, graded matched source and physical traces against the same property, and used independently qualified physical witnesses. Missing or unqualified evidence was recorded as abstention rather than agreement; the audit reported property-conditioned verdict fidelity and false-clearance risk.
The main campaign covered nine valid calibration sessions, 586 trials and 1,070 eligible source-cleared pairs. Calibration and held-out data were separated by session and cell, and repetitions between device resets never crossed the split.
Valid sessions required at least 95% trace completeness and settled-control agreement. Missing witnesses caused abstention, and the first pilot session was invalidated by camera exposure drift.
Held-out tests favored a structured risk profile
Six predictors were frozen before held-out evaluation. SimVerity won all three initial sessions against the path-only baseline and all eight confirmatory sessions; the combined tally was 11 sessions across the two cohorts. That combined tally was descriptive because the second protocol followed knowledge of cohort-1 results. In cohort 2, an exact Wilcoxon test gave p=0.0039.
The Brier score, the study's measure of probability quality, was 0.0878 for SimVerity in cohort 2, compared with 0.1995 for path-only, 0.3166 for the simulator-verdict baseline and 0.0341 for strong lookup. The comparison put the structured profile ahead of property-blind baselines but behind the per-cell lookup under stationary control, which won seven of eight sessions.
Auditability varied with the full executable setup
Auditability varied with the complete executable configuration, not only an agent's architectural label. The unmodified Hermes harness kept every simulated trace matched; two custom agent loops matched 52% to 88%; and after a registered model–client/serving configuration change, a ReAct loop matched 100%.
The property-selective pattern was also measurable under a live planner. On a frozen 32-trial M3 manifest, all 32 traces matched with zero abstention; false clearances were 4/32 for reported state, 12/32 for observable effect and 0/32 for settled control. The authors note that bounded camera capture under monotone dimming could undercount short-boundary false clearances.
The second simulator did not add an independent check
On the tested grid, the two simulators had zero eligible disagreements across 160 anchors. Strict consensus reduced coverage by 25 percentage points and conditional false-clearance risk by 1.25 points through abstention.
The result was conditional on the frozen grid and the second simulator's coarser timing. Agreement between simulators could not establish correctness if they shared blind spots.
The evidence remains deployment-local
The primary predictor covered one held-out physical path–rung pairing and 11 sessions in two cohorts. It estimated risk rather than certifying deployment; physical evidence for live planning came from one valid planner configuration, and the tests used only whitelisted lights and isolated proxies.
No human data were collected. The authors say versioned manifests, sanitized ledgers, hashes, replay code and frozen arrays can regenerate the reported aggregates. Raw camera frames were reduced to scalar statistics and discarded.
For a physical agent, the audit's deployment-local conclusion is that a simulator pass is evidence whose validity depends on the property, read boundary, path, deployment rung and qualified witness. Missing or unqualified evidence is an abstention, not agreement.
Paper data and sources
Original title: SimVerity: When Does Simulated Agent Success Survive Physical Deployment?
Authors: Zhonghao Zhan, Yefan Zhang, Krinos Li, Hamed Haddadi
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-25
DOI: Not available
Original paper · Full text