Preprint

Temperature model leads Southeast China benchmark with flexible queries

Preprint: CSTF led the reported integer-lead comparison, while half-hour and multi-grid outputs remained diagnostic.

The clearest scorecard

A query-conditioned temperature model called CSTF had the best reported aggregate results in the study’s integer-lead comparison for Southeast China. Its reported RMSE was 1.123, MAE was 0.811, bias magnitude was 0.273 and correlation was 0.993. The reported improvements over the second-best results were 6.7%, 3.2%, 17.0% and 0.1%, respectively.

The paper’s central question is whether regional temperature forecasting can be treated as continuous spatiotemporal field evaluation. CSTF is built around a shared latent meteorological state and a query-conditioned coordinate decoder. Its stated objectives address coherence across space, time and scale.

Inside the test

For the main benchmark, each sample used three hours of nine-variable ERA5 history to predict six hourly future ERA5-Land T2M fields, meaning two-metre temperature fields. The Southeast China benchmark used a 128 × 128 base grid. The Southeast China and global-scope splits each contained 4,615 training samples, 511 validation samples and 714 test samples. Those splits were chronological over 2024–2025; validation used a 0.1 fraction, and the test subset was held out for final evaluation.

Learned baselines used the same data, input-output arrangement, normalization and evaluation protocol. The study also tested controlled CSTF variants that omitted selected objective terms. Training used 50 epochs with batch size 8, AdamW, an initial learning rate of 10−3, weight decay of 10−4, cosine annealing, gradient clipping at 1.0 and random seed 42; validation loss selected the final model before held-out testing.

Where CSTF pulled ahead

At the individual supervised leads, ARROW had slightly lower errors at the first two leads. CSTF had the best reported skill from three hours after initialization through six hours.

Queries between hours and grids

The query interface was tested at 0.5, 2, 4.5 and 6 hours. The reported regional temperature structures evolved smoothly across those requested lead times. Only the integer leads could be compared directly with ERA5-Land: because the available targets were hourly, the 0.5- and 4.5-hour outputs were forecast-only diagnostics, not direct accuracy tests.

CSTF was also decoded directly at 96 × 96, 128 × 128, 160 × 160 and 192 × 192 grids. Visual diagnostics showed the main thermal patterns preserved across those resolutions rather than produced by post-hoc resizing. The diagnostics qualitatively supported resolution-controllable query behavior, but the analysis did not report a separate physical-unit accuracy comparison for every queried grid.

What the extra tests show

In the scale-consistency ablation, mean normalized cross-resolution RMSE after resampling to 128 × 128 was 1.052 for the full model, labeled “Ours,” and 1.295 for the variant labeled “w/o Lscale.” These are normalized model-space values, so the comparison is a scale-consistency diagnostic rather than a physical-unit error against ERA5-Land.

In a separate objective-term ablation, the full model’s RMSE and MAE were 1.123 and 0.811. The listed values for the variants were 1.213 and 0.858 without the auxiliary losses, 1.220 and 0.860 for w/o Lscale, 1.230 and 0.872 for w/o Lgrad, and 1.160 and 0.825 for w/o Ltemp. Across the listed variants, Ours had the lowest RMSE and MAE, while w/o Lgrad had the highest listed ablated values.

The boundaries of the evidence

The global-scope material should be read more narrowly. It was a qualitative diagnostic of large-scale thermal coherence, while quantitative comparisons remained on the native model grid. The analysis reports no confidence intervals, inferential tests or repeated-seed variability, leaving the size and stability of the differences difficult to assess from the supplied results alone.

Taken together, the study does not establish validated accuracy for non-integer lead times against direct reference targets or quantitative global forecasting skill. It also does not show that cross-resolution consistency error equals physical temperature accuracy. The clearest direct evidence remains the deterministic Southeast China benchmark; the flexible lead-time and resolution results are query diagnostics, and the global material is qualitative.

Where the record stands

The supplied document is an arXiv version-1 preprint dated 26 Aug 2026. Its appendix provides additional implementation details, extended temperature forecast case studies, and query, training, benchmark and ablation diagnostics. The authors report support from the Heavy Rainfall Research Foundation of China (BYKJ2025M14), the China Meteorological Administration Xiong’an Atmospheric Boundary Layer Key Laboratory (2025LABL-B12), the National Natural Science Foundation of China (62374031 and 62331009), and NSFC-Jiangsu Province (BK20240173).

Paper data and sources

Original title: Learning Continuous Regional Temperature Fields with Lead-Time and Resolution Queries
Authors: Chunlei Shi, Jiong Wang, Yi-Lin Wei et al.
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-26
DOI: Not available
Original paper · Full text

Versions and corrections

  1. Published automatically after legal-source, freshness, evidence, and independent-verification gates passed.