A preprint reports that two machine-learning models trained on a broad metal-organic-framework dataset were much better at predicting gas adsorption than their parent models, while keeping comparable accuracy on selected tests of bulk properties, heat capacity and neutron-scattering data.
The core dataset contains 85,524 configurations across 19,950 unique frameworks and 79 elements. It includes empty and gas-loaded structures, geometry optimizations, equations of state and finite-temperature molecular dynamics, so the training material spans both near-equilibrium structures and structures sampled in motion.
What was built
The configurations were generated with density-functional-theory calculations, or DFT, using a 680-electronvolt plane-wave cutoff. Spin-polarized calculations were used only for open-shell-metal-based MOFs. Two architecturally distinct foundation models were fine-tuned on the data. The setup also tested an explicit long-range electrostatic term beyond dispersion correction and higher-level theory.
A warning about off-equilibrium data
To see whether the moving, finite-temperature examples mattered, the study ran a training ablation, removing one slice of data while holding the validation and test sets fixed. Without 1,146 Cu-MOF molecular-dynamics frames, representing 1.7% of training, first-stage force accuracy was unchanged and energy changed by only +1.2 meV per atom. The later force-weighted training stage was approximately 25% worse.
The same comparison raised a transfer question. On chemically distinct QMOF Gaussian-noise structures, which made up 70% of validation, the no-Cu-MOF-MD run showed a +25% force-error difference, while near-equilibrium categories barely changed. Because the ablation was nonrandomized and used fixed splits in one fine-tuning recipe, the result cannot establish that the removed frames caused the difference or that a 1.7% molecular-dynamics share is universally optimal.
Comparable on standard checks
On selected Tier-1 checks, the two models remained in a comparable range. uMOF-MH had a bulk-modulus mean absolute error, or average absolute difference, of 3.92 GPa, a heat-capacity error of 8.59% and an INS Wasserstein distance, a measure of mismatch in neutron-scattering profiles, of 3.18 meV. uMOF-POLAR recorded 4.94 GPa, 7.59% and 3.27 meV on those measures, respectively.
In a separate check on two representative frameworks, the r2 SCAN-D4 reference calculations produced geometry errors 2% to 4% lower than PBE-D3(BJ). That is useful context for the labels, but it is not a broad test across the full dataset.
The clearest gains came in gas adsorption
The clearest gain appeared in adsorption enthalpy, a measure of how strongly a gas binds to a framework. The uMOF models brought the error against experiment down to a mean absolute error of 4 kJ/mol, more than halving the corresponding foundation-model error, and were reported to reproduce experimental values within the stated experimental uncertainty. The calculation used Widom insertion at 298 K for 20,000 steps.
The models also outperformed domain-specialized gas-capture models by more than 80% on adsorption-enthalpy MAE. That result came despite the reported scale of UMA-ODAC25, which was trained on roughly 70 million MOF-adsorbate configurations, nearly three orders of magnitude larger than the uMOF training corpus. It is a strong benchmark result, but it does not establish superiority across every gas, material, condition or model.
In a separate Mg-MOF-74 test, the uMOF models followed the experimental carbon-dioxide isotherm at 298 K, including the high-pressure plateau, with a slight shift in where the plateau appeared. Baseline models had qualitative shape errors: the dispersion-corrected foundation model overshot, while the uncorrected model saturated prematurely.
Binding-site identification was less consistent. In the M-MOF-74 series, uMOF-MH selected the experimentally known open-metal site as most stable for 4 of 7 metals. When it selected the correct site, the average distance deviation was 0.06 angstroms and the angle deviation was 4.5 degrees. The result was limited to that selected series, and the model did not reproduce the experimental trend for Zn-MOF-74.
A benchmark others can inspect
Alongside the models, the authors assembled a literature benchmark. Its pipeline processed 626 candidate papers into 3,986 verified property records, including 3,146 experimental records, linked to more than 650 CIF structures. The paper describes its reproducibility score as a proxy rather than a guarantee of correctness, since the benchmark depends on automated extraction and verification.
The dataset, benchmark, model weights, training and evaluation code, and benchmark pipeline were publicly released through figshare and as an ml-peg module. Those releases make the central materials available for reuse and independent checking.
The evidence has clear boundaries. Tier-1 results came from selected tests, adsorption findings from selected MOFs and conditions, and the dynamics comparison from a nonrandomized ablation. No confidence intervals or statistical significance tests are reported, and the comparisons use selected literature references rather than direct experimental validation of newly predicted properties. The preprint therefore supports a practical direction for MOF model development. Its pattern is consistent with a role for physical diversity in gas-adsorption work, but it does not show universal superiority across all MOFs, temperatures, adsorbates, properties or architectures.
Paper data and sources
Original title: uMOF: A Universal Database, Benchmark, and Machine Learning Interatomic Potentials for Metal-Organic Frameworks
Authors: Théo Jaffrelot Inizan, Prathami Divakar Kamath, Alin Marin Elena, Kristin A. Persson
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-28
DOI: Not available
Original paper · Full text