Two embodied-agent methods reached almost the same success rate in a simulated household-task benchmark, yet their recovery costs were far apart. Reflexion succeeded in 56.3% of tasks and recorded a recovery cost of 57.3, while CoPAL succeeded in 55.8% and recorded 11.9, a 4.8-fold difference.
The result comes from an arXiv preprint that treats resilience as more than whether a task is completed. The proposed framework has three parts: Rebound, which measures recovery burden; Stability, which measures instability under semantic perturbations, or changes in instructions; and Graceful Extensibility, which measures stress-related degradation. An evaluation layer is proposed for diagnosis and optimization.
A benchmark built around disruption
Researchers evaluated 400 household tasks with 10 representative embodied-agent methods in Habitat-Sim, a simulated testbed.
The study examines whether resilience can be defined, measured, benchmarked and used to guide optimization under perturbations and iterative updates. Its three metric families are meant to distinguish recovery burden, semantic-perturbation instability and stress-related degradation from success rate, safety and task completion.
The testbed used controlled perturbations affecting physical states, semantic instructions and object states, and varied stress severity.
What changed under pressure
Recovery Cost, or Crec, was 11.6 in clean episodes and 36.3 on average in perturbed episodes. The reported distributional shift was statistically significant on a Mann–Whitney test, with p below 0.001.
Stability captured a different weakness. Under semantic perturbations, overall instability β was 0.32, mainly driven by action-generation instability βout of 0.42. In plain language, changes to instructions were associated with changes in the action generated by the system, a signal the framework keeps separate from recovery burden and task completion.
Stress severity λ ranged from 0.2 to 1.0. As severity increased, relative margin decreased while recovery cost rose, with Spearman ρ = 0.82. The framework places this stress-related pattern under Graceful Extensibility.
The benchmark showed trade-offs
No evaluated method dominated all of the resilience dimensions. CoPAL had a recovery cost of 11.9 and stress capacity λ∗ of 0.85; InnerMono had stability β of 0.149; and AgentEvolver had a recovery cost of 19.2 and stress capacity of 0.21.
The wider benchmark showed different strengths in recovery cost, stability and stress capacity. In the Reflexion-CoPAL example, the reported figures concerned success rate and recovery cost: their success rates were similar, while their recovery costs differed sharply.
Optimization brought mixed results
The study then compared original and optimized variants under identical evaluation conditions, with targeted optimizations for Rebound, Stability and Graceful Extensibility.
For Rebound, the optimized variant had a recovery cost of 33.73, compared with 59.11 for its baseline, and a recovery window of 2.38 rather than 3.13. For Stability, the step-level sensitivity score βstep was 0.2552 versus 0.3185, while overall β was 0.6772 versus 0.6788.
Graceful Extensibility went in another direction. The optimized variant had a recovery cost of 99.59, compared with 36.70 for the baseline, while GE completion was 0.917 versus 0.833. Formal GE changed from Incomplete to Complete in the reported comparison, so the higher GE completion appeared alongside the higher recovery cost.
The result is bounded by simulation
The evidence is limited to the simulated Habitat-Sim benchmark: 400 household tasks and 10 representative methods. The reported comparisons therefore describe the tested setting.
Within that setting, the proposed evaluation layer is intended to diagnose process differences and guide optimization. The benchmark did not identify a universally best method; it reported different strengths across recovery cost, stability and stress capacity.
The supplied document is an arXiv preprint, version 1, dated 24 August 2026. Funding information is not reported in the supplied document.
Paper data and sources
Original title: Resilience Matters for Embodied Agents System: New Metrics, Systematic Evaluation, and Optimization
Authors: Yapeng Liu, Yuanzhao Zhai, Xudong Gong et al.
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-24
DOI: Not available
Original paper · Full text