Preprint

Preprint tests spatial AI for satellite temperature and humidity profiles

The study reports lower error variability than a pixel-by-pixel baseline, with the largest reported spatial-context difference associated with retrievals below cloud tops.

The central result is a comparison between two kinds of model. A Residual U-Net, which uses surrounding satellite pixels as well as the pixel at a target location, retrieved three-dimensional tropospheric temperature and humidity profiles from FCI observations without forecast profiles as input. Against radiosonde observations, it showed lower reported error variability than a pixel-wise 1x1-Net, with the largest difference associated with retrievals below cloud tops. Here, error variability means the spread of the residuals, or differences between the retrieval and the radiosonde measurement.

The finding concerns model development and evaluation, not operational reliability. The reported comparisons also do not include a comprehensive measure of predictive uncertainty, and they do not establish how the model would perform beyond the data and evaluation setting represented in the study.

A model that sees its neighbors

The system used all 16 FCI channels and was trained on collocated satellite observations paired with CERRA profiles over a 14-month record. No forecast profiles were supplied to it. The paper's forecast-profile-independent setup should not be read as independence from every ancillary input: the supplied analysis notes that reanalysis-derived surface pressure remained an ancillary input.

For the main rolling time split, researchers used approximately 6,048 scenes for training and approximately 1,944 scenes for validation and another 1,944 for testing. The main test set contained 11,840 valid collocated radiosonde profiles. They summarized performance with bias, the average offset from radiosonde observations, and standard deviation of the residuals. Comparisons included CERRA, the U-Net, the 1x1-Net, an ancillary-only U-Net and ERA5 climatology.

Training minimized mean squared error on standardized temperature and specific-humidity profiles. Each step used 32 randomly sampled patches of 128 by 128 pixels. Training stopped at epoch 146, and the reported model used weights saved at epoch 126 after validation loss stopped improving.

The main comparison

Across all-sky conditions, the U-Net's temperature residual standard deviation ranged from 1.5 to 1.9 K, and bias stayed below 0.5 K. The corresponding variability range was 2.2 to 3.2 K for the 1x1-Net, compared with 1.0 to 1.4 K for CERRA. On this measure, the spatial model was closer to CERRA than to the pixel-wise baseline, but its reported range did not reach CERRA's.

Relative humidity showed a similar ranking in reported variability: 12 to 20% for the U-Net, 14.5 to 23.3% for the 1x1-Net and 9 to 19% for CERRA. A positive relative-humidity bias shared across products reached about 24% at the highest levels. That shared pattern may reflect the radiosonde or reanalysis reference as well as the retrieval itself.

Cloud and daylight results need caution

The reported cloud-condition change was modest by the standard-deviation measures. Beneath cloud tops, the increases were below 0.4 K for temperature and 3% for relative humidity. In the U-Net versus 1x1-Net ablation comparison, the largest spatial-context gains were associated with below-cloud retrievals, where direct radiative information is blocked. That comparison is diagnostic, so it does not by itself establish that spatial context caused the difference.

Daytime retrievals had lower reported variability than nighttime retrievals: temperature standard deviations were lower by 0.5 to 1.2 K and relative-humidity standard deviations by 1 to 5%. Temperature bias stayed below 0.5 K in both conditions. The result bears on the study's question about whether visible and near-infrared channels contribute useful daytime information, but the comparison is confounded by seasonal composition and differences in CERRA performance between day and night. It therefore does not isolate illumination as the cause.

The winter test was weaker near the surface

On a second DJF 2025/26 test period, skill remained comparable across pressure levels from 400 to 900 hPa. Near the surface, temperature standard deviations rose by 0.2 to 0.9 K and specific-humidity standard deviations by 0.01 to 0.12 g/kg. The comparison is complicated by the use of ERA5 rather than CERRA surface pressure in that second period.

At horizontal scales of roughly 17 to 20 km, the retrieval matched and often exceeded CERRA variability. Without independent higher-resolution profile observations, the study could not determine whether that extra variation represented atmospheric structure or a retrieval artifact.

Fine-scale detail remains unresolved

Feature-importance analysis found the strongest sensitivity to the 12.3 and 13.3 micrometre channels through the atmospheric column. The 0.9- and 0.8-micrometre pair and the 3.8-micrometre band also made notable contributions. The analysis used channel substitution, however, so correlated channels can make individual contributions look smaller than they are, and the test does not measure implicit background structure learned by the model.

Several limits narrow what can be concluded. CERRA fields were used as training targets, so systematic structure or bias in that reference can carry into the retrieval. The collocated record covered 14 months, with uneven seasonal sampling, and the second test used a different surface-pressure source, making direct period-to-period comparisons harder.

The evidence remains tied to the reported FCI and radiosonde evaluation setting. The evaluation reported bias and residual spread, but comprehensive predictive uncertainty was not characterized. The study also did not have independent higher-resolution profiles to validate the extra small-scale variation.

Preprint status

The manuscript is an arXiv preprint. Its front matter says it was submitted for publication in the Journal of Geophysical Research: Atmospheres, while the supplied record does not report peer review or acceptance.

The work was funded by the Swiss Innovation Agency Innosuisse. The authors state that FCI Level 1c observations are available through the EUMETSAT Data Store.

Paper data and sources

Original title: Tropospheric temperature and humidity profile retrieval from Meteosat Flexible Combined Imager based on deep learning
Authors: Alejandro Salgueiro, Johannes Rausch, Julie Thérèse Villinger, Angela Meyer
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-26
DOI: Not available
Original paper · Full text

Versions and corrections

  1. Published automatically after legal-source, freshness, evidence, and independent-verification gates passed.