A molecular artificial-intelligence model transferred from one odor-labeling task to several others, according to a new preprint, but it struggled with the more demanding question of which mirror-image molecule should receive the higher prediction for a given odor label.
The result exposes a gap in machine olfaction: the model could make enantiomer pairs look different without consistently identifying which member should rank higher.
Strongest gains came on the main odor-label test
The researchers fine-tuned Uni-Mol2, a molecular foundation model, on the GS-LF dataset, which the paper describes as an expert-annotated collection of 4,983 molecules carrying 138 odor descriptors. The model was trained to predict multiple descriptors for each molecule with focal loss.
On molecules held back from that dataset, Uni-Mol2 recorded the best reported score on each of five summary measures. Its macro AUROC was 0.8990. The corresponding scores were 0.4077 for AUPRC, 0.3991 for F1, 0.4288 for precision and 0.4291 for recall.
The final predictions came from an ensemble of 50 models selected from the 10 best-performing hyperparameter configurations. The authors also note that their search covered 50 trials, compared with 500 in the earlier POM setup, a difference that complicates direct comparisons between systems.
Transfer held up, with trade-offs
The model was then tested directly on a Zhang dataset rather than retrained for that benchmark. The non-overlapping test set contained 837 molecules, and 108 of the dataset’s 118 odor descriptors were shared with GS-LF. Uni-Mol2 exceeded the reproduced OpenPOM baseline on all five reported measures: AUROC was 0.8975 versus 0.8937, while F1 was 0.2696 versus 0.2453 and recall was 0.4510 versus 0.4080.
The picture was more mixed on the Mayhew odor-detection task. On its 1,716-molecule non-overlapping subset, Uni-Mol2 had higher AUROC than OpenPOM, 0.8327 versus 0.8204, and substantially higher F1 and recall, 0.4417 versus 0.3187 and 0.4191 versus 0.2277. OpenPOM, however, had higher precision, 0.5308 versus 0.4669, and higher AUPRC, 0.4638 versus 0.4191.
For mixtures of odorants, the researchers used Uni-Mol2’s molecular embeddings as inputs to lightweight classical regressors, so this part was not a no-training prediction. Across pooled held-out results from the Bushdid, Snitz 1, Snitz 2 and Ravia datasets, the best Uni-Mol2 model had RMSE 0.128 and Pearson correlation 0.578, compared with 0.131 and 0.557 for the best OpenPOM model. Bushdid was the exception when the datasets were considered individually.
Sensitivity to mirror images was not enough
The sharpest caution came from a curated test of 11 enantiomeric pairs, or mirror-image forms of the same molecule. OpenPOM gave identical predictions to both members of every pair. Uni-Mol2 produced distinct predictions for every pair, with a mean within-pair L1 difference of 0.36; a paired Wilcoxon test gave p < 0.001.
But distinct predictions did not consistently point in the right direction. Among 34 odor labels that differed between paired molecules, Uni-Mol2 ranked the correct enantiomer higher in 19 cases, or 55.9 percent—only slightly above chance. Its aggregate point estimates on the 11-pair set were higher than OpenPOM’s across the reported measures, including AUROC 0.8264, AUPRC 0.5269, F1 0.0986, precision 0.0746 and recall 0.1616.
The paper’s bracketed uncertainty ranges for the enantiomer results also contain apparent inconsistencies with their point estimates. The issue affects reported ranges for AUPRC and recall for both systems, making those intervals difficult to interpret without correction.
What the model appears to have learned
An analysis of 9,453 pairs of odor labels found that labels were confused more often when they co-occurred in the training data. The Spearman correlation was 0.596, and the 10 most confusing pairs averaged 268.8 co-occurrences, compared with 5.74 across all pairs. A language-model measure of semantic similarity offered weaker evidence: the average was about the 59th percentile, and a permutation test was not significant at p = 0.16.
The authors interpret the results as support for training a model once on a canonical olfactory task and reusing its representations across tasks. They argue that three-dimensional molecular information is needed to make enantiomers look different to the model, but that three-dimensional representation alone does not solve the harder problem of mapping that difference to human-relevant odor perception.
A benchmark result, not a universal smell model
The analysis covered GS-LF, Zhang, Mayhew, a curated 11-pair enantiomer subset and four mixture datasets. The Zhang and Mayhew evaluations included non-overlapping test sets, while the stereochemical analysis relied on only 11 pairs, making that result especially tentative.
The findings support transfer on the tested datasets and tasks, not reliable prediction for every molecule, odor vocabulary or mixture. They also do not show a causal effect on human odor perception, a clinical or behavioral benefit, or zero-shot prediction of odor descriptors the model has not encountered.
The document is an arXiv version 1 preprint dated 26 August 2026. The authors say that code, datasets and supplementary information are publicly available, report no conflicts of interest, and acknowledge two University of Michigan LSA Technology Services grants awarded in 2023 and 2025.
Paper data and sources
Original title: A General-Purpose Molecular Foundation Model Transfers Across Diverse Olfactory Tasks
Authors: Yikun Han, Yi Wang, Neil Mankodi et al.
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-26
DOI: Not available
Original paper · Full text