An automated geospatial AI system identified 15 of 18 health zones that were newly invaded in five weekly forecasts in the 2026 DRC Bundibugyo Ebola benchmark. Its Recall@10 score—the share of newly invaded zones found among the 10 highest-ranked predictions—was 83.3%, compared with about 73% for the public baseline, a difference of 10.3 percentage points.
That result was part of a broader benchmark testing whether PPE could handle spatial regression, super-resolution downscaling—estimating finer-area patterns from larger-area data—and epidemiological nowcasting. The tests covered public-health, environmental-vulnerability and epidemiology domains in both Global North and Global South geographies.
PPE is organized into three modular stages: Intelligent Data Selection, Multimodal Dataset Curation, and Automated Model Building and Prediction. The stages are guided by large-language-model orchestrators.
A system built to choose its inputs
The comparison had four tiers: manual or standalone baselines; PPE with covariates, or input variables; PPE with foundation-model embeddings, or machine-generated numerical representations; and the full PPE stack. Model search spanned four supervised families and used 80:20 random splits, spatial group splits or three-fold cross-validation, with a multi-layered overfitting guard.
To reduce leakage—the risk that downstream or future information enters model development—the preprocessing filtered downstream and future covariates and calculated imputation statistics only from the training partition.
Across very different maps
The datasets operated at very different scales. The benchmark included 519 DRC health zones; 30 Nigerian ADM1 states for training and 581 ADM2 local government areas for evaluation; and approximately 84,000 US census tracts for CDC and FEMA tests under an 80:20 split. The SVI work used five county-level vulnerability scores across approximately 3,000 counties.
Where the scores rose
In Nigeria, full PPE's R²—a score for how much of the observed variation is captured by a model—was 66.1% for Food Consumption Group downscaling, compared with 31.5% for the baseline. In ADM2 out-of-fold validation, mean absolute error, the average size of a prediction miss, was 10.0% for PPE versus 13.6% for the baseline.
For US SVI county-to-zipcode downscaling, full PPE reached a mean R² of 37.6%, versus 11.0% for the baseline. The reported 95% confidence intervals, or ranges around the estimates, were 36.7% to 38.6% for PPE and 10.0% to 12.3% for the baseline.
Across 21 CDC health variables, full PPE's mean R² was 76.8%, with a reported 95% interval of 76.1% to 77.6%, versus 60% for a manual expert pipeline. Across 20 FEMA labels, full Intelligent Data Selection reached a nationwide mean R² of 64.9%, compared with 59.9% for the hand-curated model.
One result is harder to compare directly because the preprint describes its baseline in two ways. The narrative compares a peak full-pipeline R² of 66.2% with 58.6% for standalone PDFM; elsewhere, the tabulated result labels the baseline as Foundation Model Signals and gives it an R² of 60.3%, with a 55.2% to 64.9% interval, against 61.6% to 70.4% for the full stack. The baseline definition is therefore not reported consistently.
Across the benchmark suite, the paper summarizes relative R² improvements of 12% to 94% over standard baselines. It describes the lower end as a 12% SVI spatial-regression gain and the upper end as a near-doubling on Nigerian food-security downscaling, alongside the 10.3-percentage-point gain in nowcasting Recall@10.
What the benchmark can tell us
These are comparative benchmark scores from separate tasks and datasets, not one pooled performance figure. The reported evaluations mix random 80:20 splits, spatial group splits and three-fold cross-validation, so the numbers describe performance under the setups used for each test.
That scope matters: the study evaluates predictive performance in named geographies and domains. Its findings therefore support claims about model scores on the tested benchmarks, not a general guarantee that the system will perform the same way elsewhere or change health, food-security and disaster-risk outcomes.
The document is an arXiv v1 preprint dated 26 August 2026. It states that all non-confidential data sources for the Ebola hotspot analysis are publicly available through an open-access GitHub repository.
Paper data and sources
Original title: Planetary Prediction Engine: Autonomous Geospatial Prediction via Intelligent Data Selection and Foundation Model Embeddings
Authors: Evelyn Ma, Rama Kumar Pasumarthi, Kishwar Shafin et al.
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-26
DOI: Not available
Original paper · Full text