An adaptive method for blending forecasts led the aggregate scores in three online datasets, but it did not sweep the field. Gibbs-family methods recorded the best aggregate losses in experiments using Traffic Hourly, Electricity Hourly and Solar Weekly data, and won several M4 frequency and disagreement regimes. Yet familiar methods such as the median, trimmed mean, top-3 average and inverse-loss weighting still led some M4 groups. The result is a competitive forecasting tool whose performance varied with the test setting.
The central question was whether Gibbs-style weighting could improve forecast combination when models disagree by different amounts and when predictions are updated in sequence. In practical terms, the system takes forecasts from several models and combines them by assigning different weights to those model outputs.
How the method chooses its weights
The framework treats forecasting models as experts. It turns normalized predictive loss into Gibbs-style exponential weights, then adds numerical stabilization, diversity-aware score corrections and online adaptation of its hyperparameters, the settings that control the update. The approach is designed to use forecast performance while also accounting for how similar or different candidate forecasts are.
Three versions share the same core algorithm: Stable Gibbs, Directional Gibbs-NCL and Symmetric Gibbs-NCL. A companion Local-UCB procedure chooses a small set of candidate states at each forecasting case instead of evaluating the full parameter-state space every time. It adapts the learning rate, diversity strength and Gibbs variant, making online adaptation cheaper while still allowing the selected state to change.
The M4 results were mixed
For the M4 experiment, the researchers used 41 valid official numbered point-forecast submissions, treating each submitted forecast column as a base expert. For every frequency and disagreement-regime combination, they used 4,500 series, split into three independent blocks of 1,500. Each block included 500 warm-up cases and 1,000 evaluated cases, producing 3,000 evaluated cases for each combination.
The results shifted with the regime. Symmetric Gibbs ranked first in the Yearly medium, Quarterly medium and Monthly high groups. Stable Gibbs-NCL ranked first in Quarterly low and Monthly low, while Stable Gibbs appeared among the top three in several regimes. No single Gibbs version dominated the entire M4 exercise.
Conventional combination rules also led some groups. In high-disagreement Quarterly data, the median and trimmed mean were favored; in low-disagreement Yearly data, top-3 averaging and inverse-loss weighting were favored. The findings therefore point to a contest among methods, not a universal replacement of existing rules.
The M4 rankings are descriptive. The study reports no uncertainty intervals or inferential comparisons, so the rank order should be read as a reported benchmark result rather than a formal significance finding.
Online tests outside M4
Traffic Hourly's top two methods had very similar reported overall losses. Local-UCB Stable Gibbs-NCL ranked first at 421.1985, while Local-UCB Stable Gibbs ranked second at 421.2012. They evaluated an average of 4.05 and 3.00 states per forecasting case, respectively. No uncertainty interval or significance comparison was reported for this ranking.
Electricity Hourly ranked Local-UCB Stable Gibbs first, with an overall loss of 19,561,022,006 and 3.00 mean states evaluated per case. Local-UCB Symmetric Gibbs ranked second at 19,595,345,221, with 4.02 mean states evaluated. The reported analysis included no uncertainty interval or significance comparison.
Solar Weekly also ranked Local-UCB Stable Gibbs first, at 17,890,727,069. Local-UCB Symmetric Gibbs was second at 17,964,395,448, and Local-UCB Stable Gibbs-NCL was third at 18,005,974,186. The first method averaged 3.00 states evaluated per case and the next two averaged 4.36. No uncertainty interval or significance comparison was reported.
A different forecast pool
The Monash experiments used 100 Traffic Hourly sensors, 100 Electricity Hourly clients and 137 Solar Weekly series. Traffic and Electricity each started with 90 days of history and used a 270-day deployment, with 48-hour and 24-hour forecast horizons, respectively. Solar started with 30 weeks of history, rolled forward weekly and used a five-week horizon. The evaluations covered 13,000, 26,500 and 1,781 cases, with 8, 8 and 9 base models, respectively.
The source of the forecast pool also differed. M4 supplied official submitted forecasts, while the Monash datasets lacked ready-made multi-method forecast matrices in the required format, so the study generated reproducible deterministic rolling baseline forecasts before applying the ensemble methods. The two sets of results therefore come from different forecast-pool constructions.
A promising addition, not a universal winner
The evidence supports a measured conclusion. Gibbs-family methods were frequently competitive and sometimes leading, but they did not dominate every M4 regime. The method also depends on the quality, diversity and stability of the base-forecast pool, and adaptive weighting cannot recover information absent from the candidate forecasts.
The manuscript is an arXiv preprint dated 28 Aug 2026. Its evidence comes from the reported forecast pools, stream designs and deployment experiments, so it is not a universal verdict on which combination rule should lead in every setting.
Paper data and sources
Original title: Generalized Gibbs Ensemble Weighting for Forecast Combination
Authors: Prasen R. Nuthanakaluva, Nava K. Gaddam
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-28
DOI: Not available
Original paper · Full text