Some AI weather models can score well while reproducing less realistic error growth. A modeling study finds that systems can produce strong large-scale forecasts while missing the way tiny disturbances rapidly grow, a pattern associated in the analysis with the butterfly effect. In tests using ERA5 and PlaSim, physics-based numerical models reproduced that behavior, but the AI weather prediction models did not. In an idealized comparison, the model with stronger forecast skill also showed much slower large-scale perturbation growth.
The comparison was designed to isolate time direction.
The study’s core hierarchy spans ERA5 reanalysis, an intermediate-complexity general circulation model, and the multi-scale Lorenz 96 system. The authors compared models trained on coarser or large-scale-only information with finer-scale or perfect-data models and with physics-based numerical integrations.
To test whether the direction of time changed the result, forecast and backcast networks were trained independently but with matched architecture, data, optimizer and loss. Their only designed difference was whether they mapped inputs forward or backward in time. Skill was assessed with ACC, a large-scale measure; an ACC above 0.6 marked the conventional skill threshold, and the ERA5/PlaSim asymmetry was the forecast-to-backcast lead-time ratio at that point.
For the ERA5 Transformer, training used 6-hourly data from 1979 through 2018, with 2019 reserved for validation and checkpoint selection. Test initial conditions were taken once a month in 2020 and 2021, on a 1-degree-by-1-degree grid.
In the Lorenz 96 evaluation, the researchers used 100 initial conditions from the test trajectory, spaced 1,000 model time increments apart.
Forecasting worked better in one direction.
On ERA5, the forward forecast stayed above the ACC threshold for 9.1 days. The backcast—predicting backward from a later state—stayed skillful for 6.5 days, producing an asymmetry greater than 1. In plain terms, the same kind of network had a longer useful horizon going forward than backward.
The imbalance was not confined to reanalysis. PlaSim showed even larger forecast–backcast asymmetries, including in generative and neural-operator models. The pattern was also robust across variables, pressure levels and skill metrics, according to the analysis.
The missing signal was in small-error growth.
The study then asked whether these models could reproduce the rapid growth of very small errors that numerical models show. That pattern is described here as a butterfly effect: a tiny starting difference quickly becomes a much larger one. ICON and PlaSim numerical models reproduced the described behavior, but no ERA5 or PlaSim AI weather model did so, whether it was forecasting or backcasting.
The clearest accuracy-versus-fidelity trade-off came from Lorenz 96. In the perfect-data regime, AI models had approximately sevenfold worse forecast skill than in the real-world regime, but approximately four times more forecast–backcast asymmetry. Their backcasts blew up, and DKE—one of the study’s perturbation-error measures—closely mimicked a butterfly effect. The real-world-regime model’s large-scale Lyapunov exponent, a measure of how quickly nearby states separate, was around eight times smaller than in the ground-truth or perfect-data model.
A separate experiment with official Pangu-Weather models added a temporal test. When the model time step was reduced from 24 hours to 1 hour, small-amplitude DKE shifted from slow, amplitude-independent growth to rapid, amplitude-dependent growth. But the skillful lead time fell from 9.3 days to 2.0 days. The analysis notes potential confounding from high-latitude instabilities and lower-precision GPU calculations; high latitudes were excluded from the main DKE average.
Accuracy is not the whole test.
The authors interpret these patterns as consistent with a shared role for training-data coarse-graining. Their proposed explanation is that omitted fast or small-scale information can damp effective large-scale error growth and let the model absorb some missing behavior implicitly. On that reading, the simplification may help large-scale forecast scores while weakening physical fidelity and realistic error growth. These are model-based interpretations from the tested hierarchy, not randomized causal estimates.
Those conclusions have narrow boundaries. The approximate sevenfold and fourfold Lorenz comparisons come from an idealized system rather than direct atmospheric observations. The 1-hour Pangu result carries potential confounding from high-latitude instabilities and lower-precision GPU calculations. The authors explicitly caution that its butterfly-like pattern is not the real atmospheric butterfly effect.
The paper’s status and disclosures are stated openly.
The document is explicitly marked as a preprint under review. It reports support from NSF grant AGS-2531264, Schmidt Sciences LLC through InMOS, and University of Chicago institutes; the authors declare no competing interests. The authors state that all trained AI models and their inference data will be made public upon publication.
Paper data and sources
Original title: Missing the Butterfly and Predicting the Past: Features or Bugs of Accurate AI Weather Models?
Authors: Pedram Hassanzadeh, Weidong Li, Y. Qiang Sun et al.
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-26
DOI: Not available
Original paper · Full text