INCEPT, an EEG foundation model designed to be reused across tasks, posted leading results in a benchmark spanning signal-level assessment, brain-state decoding and brain-health evaluation. The study reported strong results both with the model’s backbone frozen and after full fine-tuning, depending on the task.
A benchmark built around reuse
The benchmark covered three broad jobs. Signal-level evaluation used TUAB and TUAR for clinical abnormality screening and artifact recognition. Brain-state evaluation used FACED, SEED-V, PhysioNet-MI and ISRUC-S1 for affective, sensorimotor and sleep-state decoding. Brain-health evaluation used Mumtaz2016, MentalArithmetic, ADFTD and Siena in psychiatric or depression-stress, neurodegenerative and seizure-related EEG settings.
Researchers compared task-specific supervised encoders with other EEG foundation models. They used two protocols: frozen-backbone linear probing, which keeps the main representation fixed while a task readout is assessed, and full fine-tuning, which adapts the full model to the task. Scores were reported as mean plus or minus standard deviation across five runs with different random seeds.
A clear gain on artifact recognition
TUAR provided one of the clearest frozen-backbone comparisons. On the artifact-recognition task, INCEPT recorded 58.30% balanced accuracy and 48.06% Cohen’s kappa. The reported differences from the average task-specific encoder were 17.9% and 26.5%, respectively.
When the full model was fine-tuned for TUAR, the scores were 61.45% balanced accuracy, 69.32% weighted F1 and 52.62% kappa. Compared with INCEPT’s own linear-probing result, the reported gains were 5.4%, 3.6% and 9.5%.
The frozen TUAR result was also less variable across random seeds than the task-specific encoder average. Reported standard-deviation reductions were 63.5% for balanced accuracy, 64.0% for weighted F1 and 72.7% for kappa.
Results held across several brain-state tasks
On FACED, INCEPT had the strongest foundation-model result with the backbone frozen for weighted F1, at 79.60%, and kappa, at 73.35%. The reported weighted-F1 difference from the average of other frozen foundation models was 16.8%, and weighted-F1 standard deviation was 61.2% lower. After full fine-tuning, INCEPT was reported as strongest across the metrics, scoring 84.75% balanced accuracy, 84.48% weighted F1 and 82.51% kappa.
SEED-V showed another large frozen comparison. Linear probing gave INCEPT 46.91% balanced accuracy, 45.74% weighted F1 and 33.00% kappa. Those results were reported as 31.0%, 32.4% and 74.9% above the strongest task-specific encoder, while full fine-tuning kept INCEPT highest and produced a 17.6% weighted-F1 difference over the other fine-tuned foundation models.
On PhysioNet-MI, linear probing produced 58.98% balanced accuracy, 58.65% weighted F1 and 45.30% kappa. The reported differences versus other frozen foundation models were 14.4%, 13.6% and 27.6%; after fine-tuning, the scores were 63.70%, 63.84% and 51.59%, and were reported to surpass supervised baselines.
ISRUC-S1 produced 77.33% balanced accuracy, 78.45% weighted F1 and 72.58% kappa with linear probing. Full fine-tuning produced 79.88%, 81.19% and 75.87%, respectively, exceeding the supervised average by 16.4%, 15.5% and 20.5%.
Health-related tests also favored the model
On Mumtaz2016, frozen INCEPT reached 95.51% balanced accuracy, 99.47% AUROC and 99.50% AUC-PR, exceeding the strongest task-specific encoder. The average frozen foundation-model readout had an 87.6% lower reported balanced-accuracy standard deviation, and full fine-tuning did not improve INCEPT over the frozen readout.
On MentalArithmetic, frozen INCEPT led the foundation-model comparison with 86.17% AUROC, 72.34% AUC-PR and 62.08% balanced accuracy. The reported differences versus other frozen foundation models were 14.8% for AUROC and 27.0% for AUC-PR. Full fine-tuning produced 59.72% balanced accuracy, so it did not improve that measure.
What the numbers do not show
Taken together, the numbers describe comparative model performance on named EEG datasets. They do not show that INCEPT improves human diagnosis, treatment selection or patient outcomes, and they do not establish that invariance-oriented pre-training itself caused the reported differences.
The uncertainty is also narrower than a clinical reader might assume. The study reports means and standard deviations across five random-seed runs, and the supplied sections report no confidence intervals, p-values or inferential tests. Seed-to-seed variation therefore describes stability across runs, not sampling uncertainty across participants, institutions or prospective clinical data.
Whether the model generalizes to unseen institutions, devices, montages and prospective clinical data remains open. The paper directs readers to Supplementary Section D for more detailed quantitative comparisons and dataset-specific interpretations.
Paper data and sources
Original title: Taming foundation model with invariance-oriented pre-training for broad-spectrum EEG analysis across signal-level, brain-state, and brain-health tasks
Authors: Yulong Dou, Han Wu, Guo Chen et al.
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-25
DOI: Not available
Original paper · Full text