An AI forecasting system called GENIE beat two benchmark methods more often than not when predicting daily hospitalisations and deaths in simulated outbreaks. The result comes from a computer-based test only: the study did not test GENIE against real-world surveillance data.
GENIE’s clearest advantage came in forecasts of hospitalisation peaks. Two weeks before the first peak, it put the peak in the correct week in 87.6% of forecasts, compared with 23.8% for hhh4.
A model built around place
GENIE was trained and tested on 1,000 stochastic outbreak simulations across 84 MSOAs, local geographic areas. The simulations generated daily infections, hospitalisations, deaths and Rt—the effective reproduction number—for each area.
Its graph-neural-network architecture represents neighbouring MSOAs as connected and combines location-specific profiles with local time-series interactions before making predictions.
Each forecast used a 30-day history and sampled 150 possible trajectories, allowing the system to produce a probabilistic range of possible futures rather than one fixed number.
Of the 1,000 simulations, 700 were used for training, 150 for validation and 150 for testing. The testing simulations were held back for the main evaluation.
The advantage was strongest around epidemic peaks
Against Mantis, GENIE had the best reported score in 69.78% of hospitalisation comparisons and 61.86% of death comparisons. Against hhh4, it led in 58.54% and 56.91% of comparisons using CRPS, a score for forecasts at individual locations; on Energy Score, which assesses the joint forecast across all locations, it led in 78.12% and 79.02%.
The paper did not report confidence intervals or formal significance tests for these comparison percentages. They describe performance within the reported simulation experiment, rather than establishing real-world superiority.
For the first hospitalisation peak, GENIE’s correct-week accuracy exceeded 50% when forecasts were made 17 days ahead. Two weeks ahead, it reached 87.6%, versus 23.8% for hhh4; on the second and third peak days, its lead over hhh4 was 24.3 and 26.5 percentage points.
Peak size estimates also showed lower median relative errors in the reported comparisons. For the first peak, GENIE’s error was 1.44 seven weeks ahead, 0.32 two weeks ahead and 0.05 on the peak day. Two weeks ahead, the corresponding errors for the second and third peaks were 0.19 and 0.17, compared with 0.86 and 0.78 for hhh4.
That pattern came with uncertainty: for the second peak, GENIE’s 95% error interval was wider, even though its 50% interval performed better.
GENIE was not better at every point in the outbreak. Mantis performed better during the initial 12 days of the emergence phase, while GENIE was better on most of the later initiation dates tested.
A promising result that remains untested in the real world
Researchers also removed parts of the architecture. By Energy Score, the full GENIE model was best in 76.00% of hospitalisation comparisons and 79.31% of death comparisons against a version without the Local Profile Encoder; it was best in 71.13% and 65.60% against a version using a multilayer perceptron prediction module.
The comparisons favored the complete model in the synthetic test, but they do not establish that adding either feature would improve forecasts in a live epidemic.
When the model was tested on unseen geographies, the advantage narrowed. In 150 simulations from Sunderland and Darlington, GENIE outperformed hhh4 in 41.86% and 35.26% of comparisons overall. Among the selected top 10% with the highest hospitalisation burden, the figures rose to 60.54% and 54.29%.
The manuscript is a preprint on arXiv dated 20 Aug 2026. The paper reports that code for GENIE and the JUNE simulation system is available.
The study therefore remains a test of probabilistic forecasting on synthetic epidemics, not evidence that GENIE forecasts real epidemics or improves policy, intervention timing or hospital capacity decisions.
Paper data and sources
Original title: GENIE: Generative Neural Inference for Epidemics
Authors: Laura M. Guzmán-Rincón, George R. E. Bradley, Joel Kandiah et al.
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-20
DOI: Not available
Original paper · Full text