A machine-learning method designed to predict several linked physical fields at once kept their prescribed linear relationship intact and produced the lowest reported errors in two simulation tests, with its advantage widening when the Lotka–Volterra training set was smaller. The approach, called Row-CMO, combines row-wise principal component analysis—a way to compress field data—with constrained multi-output Gaussian-process regression, which predicts several outputs together.
Row-CMO is designed so the constraint is retained not only in its predictive means but also in its posterior samples, the alternative predictions used to represent uncertainty. The paper reports that this preservation holds to machine precision under the model and constraint assumptions used in the study.
The theory behind the compression
The theoretical analysis compares row-wise PCA with field-wise PCA. It says any excess reconstruction error is governed by inter-field spectral heterogeneity—the degree to which the fields’ spectral structures differ. When those structures are similar, the penalty relative to field-wise PCA is essentially negligible.
To test sensitivity to the choice of deduced output, the scalar benchmark used three deduction scenarios across 200 independent replications. Each replication had a 200-point test set and training sizes of 20, 50 or 100 points. The constrained multi-output model used one joint training phase, while each deductive configuration was trained separately. Hyperparameters were optimized by marginal log-likelihood using L-BFGS-B, with 50 random initializations for multi-output models and 30 for single-output Gaussian processes.
Performance changed with the deduction choice
The benchmark exposed a weakness in one common shortcut. When the second output, labelled f2, was reconstructed from the others, the independent-Gaussian-process approach showed an upward shift in RMSE, the standard measure of prediction error. The linear model of coregionalization, another multi-output approach, was more stable when the deduced output changed.
Across outputs and deduction scenarios, the constrained multi-output Gaussian process, or CMOGP, beat the linear model of coregionalization more often in most configurations. The reported gap widened as the training-set size increased. Predictive intervals—the ranges intended to show uncertainty—had comparable empirical coverage near their stated target level, while their length patterns largely followed the RMSE results and added no further distinction in the reported analysis.
The strongest results came with few simulations
The authors then tested the approach on a Lotka–Volterra simulator, using four output fields on a 20,000-point temporal grid. They trained on 10, 15 or 30 simulator runs and tested on the corresponding 90, 85 or 70 runs, averaging results over 10 independent replications.
At retained dimensions m=5–10, Row-CMO consistently had the lowest relative root mean square error, or RRMSE, across all three training sizes. Its advantage was larger when fewer training runs were available. The comparison also showed that column-wise PCA captured slightly more variance in the earliest modes, but the difference between the PCA strategies approached zero beyond about m=4, and the RRMSE gap disappeared around m=5–6.
A separate CFD application used a candidate input set of 1,000, selected 50 inputs for training, tested on 100 Monte Carlo inputs and averaged results over 10 training runs.
As the retained dimension increased beyond approximately five, Row-CMO achieved the lowest RRMSE on all three outputs at m=8 and showed narrower variation across training seeds. For one representative input with eight latent dimensions, prediction errors were localized in the recirculation region, with standardized relative squared error magnitudes on the order of 10−4 to 10−3.
What the tests do not settle
Taken together, the experiments point to a possible advantage when training data are scarce: Row-CMO’s edge was larger at smaller Lotka–Volterra training sizes, and its CFD results varied less across seeds. Those findings are computational comparisons, not evidence that the model will outperform alternatives in every physical setting.
The central guarantee is stated for the model and constraint assumptions used in the study. The evidence spans one scalar benchmark, one Lotka–Volterra experiment and one CFD experiment, so it does not establish performance for other simulators, field structures, sample sizes or constraint types.
The scalar benchmark found comparable empirical coverage near the nominal level for its predictive intervals. That result does not establish calibrated uncertainty for every application.
Paper data and sources
Original title: Multi-output Gaussian process prediction of physical fields under linear equality constraints
Authors: Mahamat Hamdan Nassouradine, Clément Gauchy, Pierre-Emmanuel Angeli, Sébastien da Veiga
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-26
DOI: Not available
Original paper · Full text