A preprint evaluating Brain Researcher, a researcher-governed AI harness for neuroimaging analysis, reports a large difference in basic research navigation: models selected the correct first route or tool in 93.6% of cases with the platform and 23.3% without it. But most cited evidence still failed verification, leaving a gap between finding a route through an analysis and showing that the supporting evidence could be checked.
The document is an arXiv version 1 preprint dated 20 August 2026. Its authors present Brain Researcher as a domain-specific system operating inside an existing neuroimaging computational environment and intended to preserve scientific judgment rather than replace it.
The biggest gap came at the first decision
The core tool-calling comparison paired seven models in with- and without-Brain-Researcher conditions. Each model contributed a mean over 60 task manifests, producing 420 model-item trajectories per condition. A separate evidence-citation benchmark used 50 questions. The statistical analysis treated the seven models as the paired units and used t-based 95% confidence intervals, paired t tests and exact two-sided Wilcoxon signed-rank tests.
The pattern extended beyond the first move. Capability@1, the first-choice capability-coverage score, was 94.5% with the platform versus 49.8% without it, a 44.7-point gap. Handoff score@1, a measure of handoff sufficiency, was 76.1% versus 47.4%, a 28.7-point gap; the route/tool difference was 70.2 points. The with-platform route/tool estimate had a task-clustered 95% confidence interval of 88.8% to 97.1%, and the exact Wilcoxon test returned p=.016 for each metric.
Citation grounding was much harder. In the 50-question set, verified grounding—the share of cited evidence that passed checking—was 4.6% without Brain Researcher and 22.0% with it, a 4.8-fold descriptive increase. Even in the platform condition, most evidence rows failed verification.
A routing ablation that removed direct knowledge-graph calls still selected an acceptable exact first-choice route in 362 of 420 episodes (86.2%); model-specific rates ranged from 81.7% to 90.0%.
Testing whether claims survive scrutiny
The evaluation then moved to three collaborator-led studies and two self-evolving research episodes.
One collaborator case, a NeuroMark schizophrenia audit, used the FBIRN cohort of 363 participants—181 controls and 182 patients—with 5,460 FNC edges per subject. It ran a 480-specification multiverse, checking the hypotheses across combinations of four connectivity estimators, three confound strategies, five dimensionality-reduction methods, four classifiers and two domain granularities.
None of the three NeuroMark hypotheses was uniformly supported. NM-H2 was favorable in 12 of 24 contrasts: all contrasts under Pearson or Spearman were favorable, while none under partial correlation or mutual information was. For NM-H1, the median AUC difference was −.032; 18.8% of specifications favored latent features, and 26.0% favored between-domain loading mass for NM-H3. The estimator split remained unexplained.
A review failure showed why the human layer mattered. After a server-side fault, the system fell back to a general-purpose coding agent that used a sign-blind scoring rule. Automated review missed the mismatch; a human reviewer detected it, and directionality and fallback warnings were added.
Several analyses stopped short
Another collaborator case used 138 participants from SUDMEX CONN. Across 36 specifications, all five prespecified associations involving network systemic segregation were rejected under SDMA-GLS, and an exploratory screen of 70 combinations found no effect that survived false-discovery-rate correction.
The cross-cultural meta-analysis drew on 21 published studies and 85 peak coordinates, arranged into four culture-by-relationship cells. Its proposed cross-cultural pattern in brain-region topology was blocked as exploratory: the cells had only 6–8 entries, below the recommended minimum of 17; the paradigm mix was imbalanced, and centroid shifts alone could not establish non-overlapping distributions. No settled claim remained.
Positive signals still needed confirmation
The HCP episode began with 326 participants and 116 allocated candidate slots. Of those slots, 104 parent-run candidates returned scores and 12 ended in transport failure; the matched comparison used 244 participants from 243 families.
In that matched comparison, the workflow selected during the search and then frozen for testing reached a median correlation of .332 versus .235 for the comparator, with a median Δr of .098 and a conditional one-sided p value of .006. It exceeded the comparator in all 10 Cognition splits, but a predeclared calibration-repair decision aid failed and the analysis record marked scientific acceptance as false.
The TRIBE episode began with a 48-sound panel drawn from four source collections and balanced across six categories, while screening 15 category pairs.
Across three successive non-overlapping panels of 48 items, the measured separation between speech and tools was smaller in later layers in 11 of 12 collection-by-panel comparisons, and every panel met the prespecified directional criterion.
That recurring pattern was not confirmed at the frozen endpoint built from four new-source collections. Three of the four collections were directionally concordant and aggregate ΔS was −.1980, but the Holm-adjusted p value was .39636 against an alpha of .025; scientific acceptance was false.
A process tool, not a scientific verdict
Brain Researcher is presented as a researcher-governed, domain-specific harness that preserves rather than replaces scientific judgment. In this evaluation, its clearest gains were in route selection and handoff, while citation verification remained low and later case results were qualified, blocked or unconfirmed.
The study did not run a randomized user study or directly measure runtime or researcher effort. The NeuroMark episode also showed that automated review could miss a scoring mismatch before a human caught it, so the review layer did not replace expert inspection.
Access and disclosure
The Brain Researcher system and companion agent layer are released under the MIT license on GitHub, while BR-KG and benchmark artifacts are archived at Zenodo and linked from the project site.
The paper reports that no new human- or animal-subject data or materials were generated. One author, identified as S.K., reported recent part-time employment with Meta that began after most of the reported work; the other authors reported no competing interests.
Paper data and sources
Original title: Bringing analytic rigor to agentic AI for science: The Brain Researcher platform for neuroimaging data analysis
Authors: Zijiao Chen, Nicholas Lu, Xinhui Li et al.
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-20
DOI: Not available
Original paper · Full text