The score that set the ceiling
The best-performing system in a new arXiv preprint's test of Olympiad-level physics problems got the final answer right in 33.7% of cases. Grok-4.2 ranked first, followed by Claude-Opus-4.6 at 28.1%; every other model scored below 30%.
PhysElite is designed as a bilingual, multimodal benchmark for open-ended physics reasoning. The problems include diagrams, worked derivations and final answers, so the evaluation could examine both the result and the path taken to reach it. The reported result is therefore tied to this defined collection of scientific problems.
PhysElite contains 11,586 problems in Mechanics, Electromagnetism, Thermodynamics, Optics and Modern Physics. Each problem has a diagram, a human-verified step-by-step derivation and a final answer. The source materials came from 15 first-prize-winning local secondary-school students.
Researchers tested 18 models: eight open-source and 10 closed-source. Six were run in extended-thinking modes and 12 in standard modes. Chinese and English evaluations were scored separately and then averaged. Final-answer accuracy required a reference-matching answer; the process score gave each derivation step 1.0, 0.5 or 0.0 for fully correct, partly correct or incorrect work. Three LLM judges scored the steps, and a manual check of 200 randomly sampled problems reported a mean absolute error of 0.09 on the composite scale.
Final answers told only part of the story
Final-answer accuracy and step-level scoring told different stories. Every model had a higher average process score than answer accuracy. For Claude-Opus-4.6, the process-to-answer ratio was roughly 2.1; for LLaMA-3.1-70B, it was about 3.4. In this scoring system, a model could receive credit for parts of a derivation while still missing the reference-equivalent final answer.
The machine-to-human gap was also visible in the comparison with students. Qwen3-VL-235B-A22B recorded 20.0% answer accuracy, compared with 48.5% for the human baseline. The human process score was 65.2. That baseline came from 15 first-prize-winning students who solved the problems independently, so it offers a useful reference but remains a small recruited sample rather than a broad estimate of student or expert performance.
One comparison also examined extra visual guidance. When GPT-5.2 was given a process diagram, its reported answer score was 26.7% and its process score was 46.0%, increases of 2.7 and 3.4 percentage points over its schematic-only baseline. The comparison was not randomized, however, so the result does not show that the diagram itself caused the change.
Where the reasoning broke
The error analysis looked at where a derivation first went off course. Claude-Opus-4.6 first erred at step 2.48 on average, while weaker open-source models failed near step 1.4. Qwen2.5-VL-7B failed at the first step in 96.8% of cases. Across models, soft failures ranged from 7% to 46%, and failures confined to the last step reached 7.8%. The pattern suggests that stronger systems generally got further before breaking down, but it does not explain why.
The scores also varied by subject and difficulty. Optics had the most frequent setup failures and Thermodynamics the fewest. Among the top eight closed-source models, the gap between those subjects exceeded 20 percentage points. Four frontier reasoning models showed score gaps of more than 10 percentage points between Hard and Medium problems, part of a non-monotonic pattern in which Medium problems often scored below Hard ones.
Among 100 incorrect Claude-Opus-4.6 predictions, physical-perception errors accounted for 31% and physical-law errors for 25%. Algebraic or symbolic errors made up 18%, incomplete derivations 13%, geometric or visual errors 8%, and numerical errors 6%. The largest reported categories were therefore physical-perception and physical-law errors, while numerical errors were least common. This breakdown covers one model's 100 incorrect predictions, not all evaluated systems.
A benchmark with clear caveats
The benchmark comes with important qualifications. The authors acknowledge ambiguity in partial-credit process scoring and possible minor overlap between training data and commercial model datasets. They say a fully private dataset is planned to reduce leakage risks. No confidence intervals or formal significance tests are reported, so the apparent gaps cannot be treated as precise population estimates.
The dataset's construction also shapes what it can show. Images were scanned at 300 dots per inch, processed with OCR and manual correction, and converted into LaTeX by hand. Deduplication at a 0.95 similarity threshold reduced an initial pool of around 15,000 problems to 11,586. Because the source material came from local contestant training materials, its representativeness of other Olympiad populations is uncertain.
The document is listed as arXiv:2608.25097, version 1, dated 25 August 2026. The authors state that the PhysElite datasets have been released on Hugging Face. By combining bilingual problems, diagrams, final-answer scoring and step-level diagnostics, the benchmark offers a focused test of multimodal scientific reasoning, but its findings remain specific to the tasks and evaluation setup used here.
Paper data and sources
Original title: PhysElite: How Far Are LLMs from Solving Olympiad-Level Physics Problems?
Authors: Ruoran Xu, Wending Gao, Liyunfeng Chen et al.
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-25
DOI: Not available
Original paper · Full text