A reporting system that binds each statistical claim to its evidence before an AI model writes the surrounding prose was far more consistent than a hybrid template in tests of brain-imaging and clinical-trial reports. In the main fMRI comparison, 98.5% of the numbers visible in reports were reproduced across runs for claim-locked reporting, compared with 61.1% for the hybrid method. The 37.4-point gap had a 95% confidence interval, the study's uncertainty range, from 15.1 to 59.7 points.
The change came before the prose
The protocol puts the control point before the prose. It binds each reportable claim to its evidence source, numbers, direction and permitted language strength before the LLM writes connective prose. The evaluation compared seven reporting methods.
The fMRI evaluation used 428 participants from an institutional obesity-focused cohort and 712 adults from the public HCP S1200 cohort. It ran each method across two writer providers and five seeds, yielding 20 cells per method. The RCT benchmark used 200 Evidence Inference 2.0 records: 112 described significantly increased outcomes and 88 described significantly decreased outcomes.
The numbers held steady
Across the fMRI methods, grounded baselines reached 15.2% to 32.3% cross-seed reproducibility. The hybrid template reached 61.1%, while claim-locked reporting reached 98.5%. In HCP alone, the rates were 98.0% for claim-locked, 62.0% for hybrid and 46.0% for free-form generation.
The fMRI comparison also reported language-strength and numerical-audit measures. The unhedged-strong count was 0.75 for claim-locked reporting and 2.25 for the hybrid template. In plain terms, this counts sentences that use strong language without a hedge. Both methods had a 0.00 per-run numerical-audit flag.
In a blinded human audit, mean unhedged-strong violations were 0.60 per claim-locked report, 1.85 per hybrid report and 2.55 per free-form report. Mean unsupported-content violations were 0.53, 2.33 and 1.65, respectively.
The fMRI work included a separate BMI absorption stress test. After BMI was added as an adjustment variable, the number of group-effect edges that survived false-discovery-rate correction fell from 10,041 to 8 in the institutional cohort, a 99.9% drop, and from 550 to zero in the HCP comparison, a 100.0% drop. The authors describe this as a stress test, not a biological finding about obesity.
Trial directions exposed the risk
The RCT benchmark showed the same reproducibility pattern. Claim-locked reporting reached 100.0% reproducibility, compared with 79.5% for the hybrid template, a 20.5-point difference. The reported 95% confidence interval for that difference ran from 17.8 to 23.1 points. The raw numerical-audit flags were 0.23 and 0.11, respectively.
A direction audit tested whether reports preserved an already resolved increase or decrease label. Across 120 reports per method, claim-locked reporting preserved 116 directions, inverted none and left four unresolved. The hybrid template preserved 102, inverted 17 and left one unresolved. With unresolved reports excluded from the denominator, the inversion rates were 0.0% and 14.3%, respectively.
Manual calibration of the RCT flags separated harmful statistical-result fabrication from raw alerts. Harmful fabrication appeared in 0.0% of claim-locked outputs, 2.5% of hybrid-template outputs and 5.9% of free-form outputs. The raw flags numbered 181, 89 and 213, respectively. The figures show that raw alerts and harmful fabrication were not the same measure.
Nested component tests reported 85.0% fMRI reproducibility and 90.3% RCT reproducibility for Builder-only configurations. Builder plus renderer configurations reached 98.0% and 100.0%, respectively. The full configuration reached 98.5% on fMRI, with a Strong count of 0.75 compared with 1.40 for Builder plus renderer.
A safeguard, not a statistical referee
The claim-locked setup also had the lowest observed resource measures in the reported DeepSeek fMRI comparison: 12,482 input tokens, 4,209 output tokens and 77.3 seconds of latency.
One diagnostic was less straightforward. On a 30-record RCT subset, claim-locked reporting ranked lowest on FActScore, at 0.552, but highest on SelfCheckGPT, at 0.974. The two automated diagnostics pointed in different directions.
The central caveat is that claim-locked reporting verbalizes structured evidence. It does not validate statistics, discover mechanisms or guarantee truth beyond the evidence record. Upstream errors can still propagate into the final report, and reproducibility alone does not guarantee correctness.
The RCT evaluation covered increased and decreased direction labels, but not null or inconclusive findings. Public HCP-based fMRI and Evidence Inference 2.0 settings were used, while the institutional FC cohort cannot be redistributed. The authors report releasing code, prompts, generated reports, audit scripts and permitted evidence records.
Paper data and sources
Original title: Provenance Before Prose: Claim-Locked Reporting
Authors: Xiao Fan, Jingyuan Li, Hongbin Guo et al.
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-26
DOI: Not available
Original paper · Full text