Preprint

Preprint reports lower-error approximations for quantum cloud-cover models

Fourier and interpolation-based reconstructions often showed lower mean squared error in finite-shot tests, while hardware results remained entangled with calibration effects.

A new preprint reports that trained quantum machine-learning models for cloud-cover prediction can be approximated, after training, by constructive “shadow” models that often showed lower mean squared error in finite-shot tests. The comparisons used the same shot counts for the shadow and direct quantum-model evaluations.

The study tested two ways to build those shadows. One reconstructs the model with a truncated Fourier series, retaining the largest coefficients; the other uses piecewise-affine quasi-interpolation, a grid-based approximation that does not require knowledge of the quantum circuit’s encoding strategy. Because the methods are constructive, they do not require a separate training or regression stage.

The paper is identified as arXiv version 1, dated 20 August 2026. Its results concern model approximation and prediction error under sampling and hardware conditions.

A large test set, with smaller hardware checks

The evaluation used nested test sets derived from DYAMOND cloud-cover data. The main set, D, represented around 37 percent of the parent set D0 and contained approximately 9 million cells. Resource-intensive checks used a reduced set, Dr, with 4,000 cells, while a sparse-grid evaluation used Dc, a 1,000-cell set of cirrus-cloud cases.

The main yardstick was mean squared error, or MSE: the average squared difference between a model’s prediction and the reference value. The researchers also compared the shadow model’s MSE with the direct QML model’s MSE, and examined how performance changed as they added grid points or retained more Fourier frequencies.

The comparisons covered noiseless calculations, finite-shot simulations and tests on Euro-Q-Exa hardware. The study evaluated two trained cloud-cover QML parametrizations and included state-vector simulations for comparison. The evidence therefore describes the behavior of these approximations in the tested setup, rather than establishing that they will work the same way with other encodings or hardware systems.

Interpolation helped under sampling noise, but grids grew quickly

For full-grid quasi-interpolation, the reported convergence rate was approximately O(N^-0.5), where N is the number of grid points. In practical terms, the approximation improved as the grid was refined, but the grid size needed for that improvement limited the method’s practicality.

When the models were evaluated with a finite number of shots, interpolant shadows had lower MSE than direct QML in most of the tested cases at the same shot count. The trade-off was that higher shot counts required finer grids, adding to the computational burden of the approximation.

That pattern became less consistent when the underlying QML models had been trained with 1,000 shots and variance regularization. In that setting, the interpolation result appeared only on finer grids and was weaker than in the other reported finite-shot comparisons.

The researchers also explored sparse-grid interpolation, intended to reduce the number of grid points in higher-dimensional calculations. In these tests it was comparable to, or worse than, the full-grid method in MSE and showed no improvement in convergence rate, so the approach was not pursued further.

Fourier shadows kept only part of the spectrum

The Fourier method offered a different way to control the approximation. When the spectrum is equidistant, the approach provides an exact circuit representation before truncation. The tests then examined how many of the largest coefficients could be retained while keeping the shadow close to the direct QML error.

The reported threshold was a shadow-to-QML MSE ratio below 1.05. About 7,500 of 15,625 frequencies were needed for the ZZXY model, while about 12,500 of 531,441 were needed for the XYZ model to meet that criterion. The required retained count therefore varied substantially between the two models.

For models trained with infinite-shot evaluations, the truncated Fourier shadow had lower MSE than direct QML when the Fourier coefficients were evaluated with either 100 or 1,000 shots. After 1,000-shot training with variance regularization, the same lower-MSE pattern was weaker but remained observable.

The hardware result comes with a warning

On Euro-Q-Exa, with 1,000 shots for each expectation value, the shadow model had a lower average MSE than direct QML in the reported comparison. But the study could not separate that difference from hardware calibration and other system effects, so the result is suggestive rather than a clean test of shadowing alone.

The hardware runs also showed a calibration-related pattern. Runs performed on the same day had the two smallest pairwise MSEs, although the reported difference was not statistically significant. Discrepancies between hardware runs were an order of magnitude smaller than the differences between hardware results and state-vector simulations that included simulated finite-sampling noise.

A possible shortcut, not a quantum advantage

The authors interpret Fourier truncation and interpolation as error-mitigation strategies that could make a trained QML model easier to use classically after training. But the comparisons were between shadow models and direct QML evaluations; they did not establish an advantage over a competitive independent classical machine-learning model.

The evidence is limited to one cloud-cover application and two QML architectures, mainly using six input features, with exploratory tests of a smaller model on Euro-Q-Exa. The reported outcome was offline MSE. The study did not test whether the shadow models improve online climate simulations, preserve their stability or improve end-to-end climate-model behavior.

The supplied analysis reports no formal uncertainty intervals or inferential tests for the MSE comparisons. The reported truncation thresholds were selected experimentally, and their transfer to other datasets, encodings, hardware systems or generalization settings remains unresolved.

The open questions include whether adaptive constructive methods can scale to higher-dimensional and non-equidistant encodings, whether hardware-aware training transfers across calibration regimes, and whether the observed pattern survives metrics other than MSE.

Supporting source code, Euro-Q-Exa data and test data were reported as hosted through GitHub and Zenodo. The acknowledgments report institutional funding and computing support from the DLR Quantum Computing Initiative, the Federal Ministry, the DFG and DKRZ resources.

Paper data and sources

Original title: Shadow models of a quantum model for cloud cover and the influence of finite sampling noise
Authors: Hedwig Keller, Mierk Schwabe, Veronika Eyring
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-20
DOI: Not available
Original paper · Full text

Versions and corrections

  1. Published automatically after legal-source, freshness, evidence, and independent-verification gates passed.