Preprint

Neural model processes 2.1 million electromagnetic soundings in 25 seconds

Preprint: Held-out tests showed close solver agreement; matched posterior tests found similar but not identical behavior, while the full-sampler comparison was limited by incomplete mixing.

A neural-operator surrogate inverted 2,131,667 in-range airborne electromagnetic soundings in 24.9 seconds on one GPU in the study's survey-scale application. The reported comparison estimated about 26,300 years for the full numerical solver, using 38.8 seconds for each forward solve. The speed result covers the in-range survey soundings used in that application.

The work is described in an arXiv v1 preprint dated 26 August 2026. Its central question was whether one surrogate could learn a three-dimensional airborne electromagnetic forward map across sequential geological priors, detect inputs outside its support and reproduce solver-based Bayesian posteriors. The posterior tests were reported on matched comparisons involving 15 soundings, not across the 2,131,667-sounding survey-scale application.

A model built to carry priors forward

The evaluation combined the 2013 TEMPEST survey, with 2,155,272 soundings on 191 flight lines, with one prior for the survey area and four priors from comparison regions. For each sounding, the system used a 32 by 32 by 24 voxel local crop. Training examples paired those inputs with a three-dimensional finite-volume OcTree solve, while the surrogate used a model encoder, a prior embedding and a query network.

Bayesian inversion used PPM proposals with tempered Metropolis acceptance to sample the conductivity posterior. Continual training used replay memory and function-space distillation to preserve information from earlier priors. A five-member ensemble used the 95th percentile of disagreement to route out-of-support soundings to the full solver.

Accuracy depended on the prior

On 125 held-out Capricorn soundings, the surrogate reached R2 of 0.992 and a median gate error of 4.7%. Across ten initializations, its error was 5.0% plus or minus 0.3 percentage points. Capricorn error continued to fall as the training set grew from 60 to 1,000 soundings, following a fitted power-law exponent of -0.46.

Without further training, median gate errors were 6.6% for Denmark, 5.4% for Wisconsin, 11.8% for Zeeland and 67.7% for Seward. The validity screen routed 32.8% of Seward soundings and 16.8% of Zeeland soundings to the solver, compared with 0.8% each for Denmark and Wisconsin. Capricorn test data had a stated routing fraction of 4.0%.

The sequence mattered

Continual training used replay memory and function-space distillation, and its final errors across four arrival orders were 4.4%, 4.9%, 4.4% and 4.5%. Fine-tuning ended at 11.6%, 18.9%, 15.4% and 25.9% in those four orders. Fine-tuning showed 7.5 percentage points of later-stage forgetting on average, while the continual model's average change was minus 0.6 percentage points.

Retraining all five priors at once reached 4.6%, close to the continual sequence's 4.4%. Across the tested arrival orders, the continual model remained more stable than fine-tuning.

Posterior agreement came with a warning

In a matched comparison involving 15 soundings from line 1005801, the surrogate and solver posteriors shared 85.5% of accepted realizations. The surrogate residual was 11.3%, versus 12.2% for the solver, with a per-depth difference of 0.45. In the full-sampler comparison, the surrogate range was 96% as wide as the solver range, the combined extent overlap was 85%, and acceptance was 0.30 for the surrogate versus 0.28 for the solver.

These figures describe similar but not identical posterior behavior in the matched tests. The full-sampler comparison was limited because neither engine was fully mixed at the reported budget: the Gelman-Rubin statistic was 1.39 for the solver and 1.47 for the surrogate.

Credible-interval calibration on synthetic truths was close to the intended levels. Coverage was 50.6% for a 50% interval, 79.3% for an 80% interval, 88.7% for a 90% interval and 92.4% for a 95% interval. The largest departure from the nominal level was 2.6 percentage points, and the smallest was 0.6.

What the survey produced

At survey scale, the surrogate inversion of 2,131,667 in-range soundings took 24.9 seconds on one GPU, or 11.7 microseconds per sounding. The reported resistivity summaries rose with depth at 84% of soundings, from a median of 63 ohm-metres at 38 metres to 340 ohm-metres at 238 metres. At depth, the north-western block was reported at 929 ohm-metres, compared with 264 ohm-metres elsewhere; posterior-width factors were 2.1 at 38 metres and 5.0 at 238 metres.

In a prior-updating experiment, a layer-correlated two-component mixture increased the number of explained unseen lines from 127 of 190 to 184 of 190, without losing any lines already explained. After the mixture was folded into the operator, error on the new prior fell from 28.1% to 8.1%; earlier accuracy and the posterior for the original line were unchanged.

Still a preprint

The Capricorn survey data are reported as publicly available from Geoscience Australia under eCat record 81642, while the code is promised on GitHub after publication acceptance.

Paper data and sources

Original title: Continually learning neural-operator surrogate for three-dimensional airborne electromagnetic Bayesian inversion
Authors: Jaehong Chung, Andrew Lockwood, Jef Caers
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-26
DOI: Not available
Original paper · Full text

Versions and corrections

  1. Published automatically after legal-source, freshness, evidence, and independent-verification gates passed.