Preprint

AI chemistry system scores highest at recovering past ideas

A preprint describes an automated framework that retrieves literature and revises hypotheses, but its chemical ideas were not tested in the lab.

An automated system for generating chemistry hypotheses recorded the highest listed overlap scores on the TOMATO-Chem benchmark. MAIL exceeded the MOOSE-Chem baseline on all four reported measures: scores for matching a target paper's main idea and main points, both at the highest single output and across the average of ten finalized outputs. The result measures how closely generated hypotheses recover reference ideas, not whether the system has produced a chemically workable discovery.

A benchmark of remembered ideas

That distinction is central to what the researchers set out to test. The study asks whether a fully automated process can use literature retrieved from different time periods, compressed memory and model-generated feedback to recover the central ideas and methods of historical chemistry hypotheses. It therefore tests alignment with known targets. It does not establish independent originality or show that a generated hypothesis would survive a wet-lab experiment.

The test covered two chemistry and materials-science benchmarks. The public TOMATO-Chem set contained 51 target papers. The authors also evaluated a newly curated HN-NS set containing 50 recent chemistry papers from Nature and Science. These papers served as the reference targets against which the generated hypotheses were compared.

A loop of retrieval and revision

MAIL is designed as a chain of linked steps. It draws inspiration from the literature, refines a draft incrementally through self-evaluation, adapts its prompts as the process develops and then selects a final hypothesis. The framework combines retrieval, compressed memory, feedback and final selection within one automated process.

For the main comparison, MAIL and the baseline implementations were run in the same experimental period with the same GPT-4o configuration and temporal cutoff. Each method was run independently ten times for each background. The generation setup used GPT-4o at temperature 0.7 and three refinement rounds. Literature was divided into pre-2015, 2015 to 2020 and 2021 to 2023 intervals. Each execution produced 8 to 10 candidates per round, or 24 to 30 candidate outputs overall, and required about 35 to 40 LLM calls.

What the numbers capture

On TOMATO-Chem, MAIL's Top MIOS was 4.27, versus 4.01 for MOOSE-Chem. Its average MIOS was 3.69, compared with 3.15. For main-point overlap, MAIL reported a Top MPOS of 3.91 and an average MPOS of 2.38, while MOOSE-Chem scored 3.56 and 2.14.

The average figures are the mean across ten finalized outputs. The Top figure is the post-hoc maximum among those same ten outputs, calculated only after generation, refinement and selection. That distinction matters because the Top result captures the strongest output in a set, while the average describes the set as a whole.

The numerical gap should still be read narrowly. The supplied analysis reports no confidence interval or inferential comparison for these table values. The result documents higher reported benchmark scores in this experiment, but it does not by itself show how reliably the difference would hold under other conditions.

The limits of the result

The authors interpret the pattern as indicating that dynamic retrieval, compressed memory, internal feedback and adaptive prompts can support coherent, mechanistically plausible and target-aligned hypotheses. They frame MAIL as a possible early-stage discovery tool, while acknowledging that its outputs remain unvalidated in wet-lab experiments. The study's evidence is therefore about reconstructing ideas from a benchmark, not independently finding and testing new chemistry.

The overlap measures are retrospective reference-recovery scores, so they do not directly measure independent originality. Generated hypotheses have not been physically or wet-lab validated. The comparisons were not designed to estimate runtime, token use, API cost or computational efficiency. Evaluation focused mainly on organic and materials chemistry, and the authors flag potential dual-use risks involving hazardous chemical pathways.

Prospective experiments, broader testing across specialized chemical fields and a clearer accounting of computational cost would be needed before the framework's usefulness outside retrospective benchmarks could be judged.

Paper data and sources

Original title: MAIL: Memory-driven, Adaptive, Incremental, and Literature-grounded Framework for Hypothesis Generation in Chemistry
Authors: Mahdi Babaei, Xueshen Li, Yutao Kuang et al.
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-28
DOI: Not available
Original paper · Full text

Versions and corrections

  1. Published automatically after legal-source, freshness, evidence, and independent-verification gates passed.