In the reported evaluations, DExtrI was the top performer on the synthetic benchmark, seven of eight HPOBench tasks and every displayed Seoul Bike Sharing Demand setting. It also led on the O’Neil drug screen, but not on ALMANAC, where the Neural Additive Model had the lowest reported error.
The study asks whether so-called axis-aligned training samples—data with one covariate active at a time—can support predictions for test inputs with several covariates active together. Here, extrapolation means extending the learned response to combinations not represented in the training support.
How the test worked
DExtrI uses an additive structure with interaction components represented by ridge functions, a mathematical form for representing interactions. The comparison included a Neural Additive Model, Empirical Risk Minimization and Engression; ERM and Engression were sized to roughly match DExtrI’s parameter count.
The training objectives also differed: NAM and ERM used mean squared error, while DExtrI and Engression minimized the negative energy score, an objective for estimating conditional distributions.
In the drug-screen experiments, single-drug concentration sweeps supplied the training data and combination experiments supplied the test data. O’Neil contributed 11,856 training observations and 358,560 test observations; ALMANAC contributed 1,363,939 and 2,045,907, respectively.
Results shifted by dataset
On O’Neil, DExtrI’s test root mean squared error—the reported measure of prediction error—was lowest at 1.045, with a standard deviation of 0.042 across seeded runs. On ALMANAC, NAM was lowest at 0.731, with a standard deviation of 0.002, while DExtrI recorded 0.820, with a standard deviation of 0.031.
In HPOBench, the main text reports 39 training observations per problem from one-at-a-time hyperparameter sweeps. Joint configurations were used for testing, and every test point lay outside the convex hull of the training support, or the region spanned by those training points.
DExtrI had the lowest test MSE on seven of the eight HPOBench tasks. Engression was the exception, performing best on the blood-transfusion task.
On simulated data with 40 covariates, the evaluation used 4,000 training points and 5,000 test points, with results averaged over 10 seeds. DExtrI had the lowest test MSE in every reported setting, frequently by one to several orders of magnitude.
A test on approximately axis-aligned Seoul Bike Sharing Demand data gave DExtrI the lowest displayed test MSE at 10%, 16% and 25% training. The corresponding results were 0.60 with a standard deviation of 0.06, 0.53 with a standard deviation of 0.03, and 0.44 with a standard deviation of 0.01.
A conditional mathematical claim
Under its stated model and identifiability conditions, the analysis says the structural components are identifiable—that is, recoverable within the defined model class—and that the true and fitted transformations agree for every input in the training or product test support, with a fixed noise value.
That is a model-based theorem, not a finite-sample validation result, and its guarantee holds only under the stated assumptions.
The empirical evidence is limited to the reported simulations, HPOBench problems, drug-screen datasets and approximately axis-aligned real-data evaluations, so the conclusions remain tied to those computational settings.
The manuscript is a preprint: its front matter identifies it as arXiv:2608.19849v1, dated 20 August 2026.
Paper data and sources
Original title: Distributional Extrapolation for Interactions
Authors: Marin Šola, Xinwei Shen, Peter Bühlmann
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-20
DOI: Not available
Original paper · Full text