Preprint

AI Physics Benchmark Finds Accuracy Falls at University Level

Preprint: OmniPhys tests multimodal models on Chinese physics questions, reasoning tasks and diagram generation.

A new preprint reports that multimodal AI models performed best on junior-high physics questions and declined at each successive educational stage toward university level. The benchmark, called OmniPhys, is introduced as a unified Chinese test of physics mastery from secondary education through university.

OmniPhys contains 15,246 questions and 19,850 images. The questions cover mechanics, electromagnetism, optics, thermodynamics and acoustics. The reported educational breakdown is 4,524 junior-high questions, 9,217 senior-high questions and 1,505 advanced university questions. Its task mix includes 10,227 objective questions, 4,258 open-ended questions and 761 multimodal-generation tasks.

A test of answers, working and visuals

The benchmark evaluates more than whether a model reaches the right answer. Its scoring framework separates objective-task answer results, labelled S1, from process quality on objective tasks, labelled S2, and process quality on open-ended tasks, labelled S3. It also reports two mastery scores: Pobj for objective tasks and Popen for open-ended tasks. In the main results, the Gemini-3-Pro row reports 69.82 for Pobj and 49.30 for Popen.

The researchers checked whether questions genuinely depended on their images. In a random sample of 300 instances, the visual-dependency classification agreed with expert judgment 89.3% of the time. Cohen's kappa was 0.82.

The questions were curated from authoritative Chinese teaching resources and authentic examination papers spanning middle school to university curricula. The researchers used semantic deduplication to remove near-redundant items, with a cosine-similarity threshold of 0.95. The process eliminated 12.1% of redundant instances.

Automated judges were more generous about diagrams

The paper's clearest caution concerns diagrams generated by the models. Automated evaluators gave scores ranging from 0.51 to 0.75, while human judgments ranged from 0.10 to 0.29. The authors interpret this gap as evidence that automated diagram judges can overrate outputs, and say human review remains necessary for rigorous physical checking.

The researchers also created a smaller set called Test-Mini by ranking samples with a hardness score and selecting the top 10%. The score assigns a weight of 0.75 to empirical difficulty and 0.25 to reasoning complexity. Test-Mini contains 1,288 samples, compared with 12,885 in the full non-generation set.

In the reported comparison, all-model failure was 15.5% in Test-Mini and 1.6% in the full set. Average model accuracy was 23.52% in Test-Mini and 72.44% in the full set. Test-Mini therefore showed more all-model failures and lower average accuracy than the full set.

What the benchmark can and cannot show

Several evaluation measures tracked human judgments closely. Human-machine alignment produced Pearson correlations of 0.89 for S1 and 0.92 for Pobj, with agreement rates of 92.4% and 93.1%, respectively. That pattern was narrower for generated diagrams, where the automated and human score ranges diverged sharply.

The benchmark's reach is limited by its focus on Chinese educational materials. Its scores do not establish how models would perform on other languages or curricula, and benchmark results alone do not show that a model has mastered physics in real-world settings.

The work is a preprint, identified as arXiv:2608.25398v1 and dated 26 Aug 2026. The authors state that the code and data are available at https://github.com/ECNU-RAIL/OmniPhys-EMNLP2026. They report funding from the Guangxi Science and Technology Program and the Open Research Fund of a statistics and data science laboratory at East China Normal University.

Paper data and sources

Original title: OmniPhys: A Unified Multimodal Benchmark for Physics Understanding and Generation from Chinese Educational Corpora
Authors: Hao Chen, Yumin Lin, Nadila Yushanjiang et al.
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-26
DOI: Not available
Original paper · Full text

Versions and corrections

  1. Published automatically after legal-source, freshness, evidence, and independent-verification gates passed.