In the benchmark, a frozen large language model recorded a much higher score for recovering the terms in simulated governing equations from a compact physical interpretation of field data than from raw field slices, according to an arXiv version 1 preprint dated 25 Aug 2026. Across 44 samples, mean F1 was 0.720 for the interpreted input and 0.225 for raw slices. F1 is the benchmark’s measure of how much of an equation’s term set was identified, while exact recovery requires the entire proposed term set to match.
The study asks whether giving a large language model a direct interpretation of field data can support discovery of a partial differential equation, or PDE, instead of relying only on evaluator feedback. Here, the target was the governing equation structure that the model was asked to recover for each simulated field.
What the model saw
The test used noise-free simulated fields and a fixed library of eight possible terms. Each field was integrated on a 64 by 64 grid over 50 frames, and the 44 samples were distributed across six benchmark strata.
The system ran through four stages: data interpretation, equation proposal, parsing and canonicalisation, and evaluation, using the frozen QwQ-32B model. Spectral analysis was the main part of the interpretation stage. The compact input used approximately 310 interpretation tokens. Controls included an interpretation permuted from another sample, ten raw one-dimensional slices using approximately 1,230 tokens, and the best fixed term set as a baseline. For each proposed structure, regularized least squares fitted the coefficients, while a weak-form residual and sparsity penalty scored the equation. The regression step did not choose the structure.
A wide gap in the main score
In the headline comparison, mean F1 was 0.720 for the interpreted input and 0.225 for raw slices. Exact recovery occurred in 14 of the 44 samples for the interpreted input, compared with 2 for raw slices. The reported raw-slice exact-recovery denominator is inconsistent between the table and the prose.
Against the best fixed-term-set floor, data interpretation exceeded it by 0.329 mean F1. The reported p-values were 0.000026 for the paired Wilcoxon comparison and 0.0066 for exact recovery. Those statistics come from the 44-sample simulated benchmark.
A separate control permuted interpretations between fields. Mean F1 was 0.244 for the permuted condition, compared with 0.720 for matched interpretations, with a reported p-value below 0.0000003. The comparison was intended to assess whether an interpretation corresponded to the field it described.
The score was not consistent across cases
Scores varied by benchmark stratum. For the interpreted input, F1 was 1.000 for linear-only samples, 0.500 for reaction, 0.579 for observable self-advection, 0.625 for suppressed self-advection, 0.818 for wave and 0.685 for weakly identifiable wave. The highest reported score was for the linear-only cases, while the lowest was for reaction cases.
The paper identifies a structural limit in the representation. The interpretation can show that a nonlinear term is present without identifying its form. The authors say this defines the space of PDEs the method can discover, and that nonlinear terms were identified less accurately.
A benchmark with a clear boundary
The reported characterizations are fragile to noise, so robust differentiation and noise filtering would be needed before experimental application. The benchmark itself used noise-free simulated fields and did not test noisy experimental field data.
The authors conclude that an LLM given small physical measurements identifies governing equations better than one given the field itself. They also report a fraction-of-a-second cost per field, fewer tokens and no training.
Those conclusions are bounded by 44 samples, one frozen QwQ-32B model, noise-free simulations and a fixed eight-term library. The preprint does not show that the approach works on noisy experimental fields, generalizes beyond those benchmark conditions, or generally recovers nonlinear term forms exactly. Open questions include whether robust noise handling preserves the result and whether richer characterization tools expand the discoverable PDE space.
Paper data and sources
Original title: What Should a Large Language Model See? Physical Invariants as a Data Representation for PDE Discovery
Authors: Fan Yang, Matt Thomson
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-25
DOI: Not available
Original paper · Full text