An arXiv preprint reports that a decoder recovered measurable information about the word being read from non-invasive EEG during silent reading. Across 576 model fits, the primary within-run retrieval gain averaged 6.7 percentage points over an empirical permutation baseline; the median was 6.3 points, and results ranged from 1.4 to 15.4 points.
The result came from one participant recorded intensively, contributing 240,141 word presentations across 393 runs and 48.7 hours of recording. It is therefore a detailed one-person result, not a population-level accuracy estimate.
How the word signal was measured
The participant read continuous fictional narrative prose one word at a time, with typography randomized from trial to trial. The recording used 19 dry scalp electrodes sampled at 600 Hz.
The researchers represented each target word using either a non-contextual or contextual layer of the Llama-3.1-8B language model. An EEG encoder, with an optional four-layer causal transformer to track sequence information, was trained to align those targets in a shared 256-dimensional space.
For evaluation, about 80% of runs were used for training and 20% for validation. The system ranked supplied candidates in fixed pools of 512 trials, with the top-10 score averaged over five pool draws against baselines made from 10 permutations. This was a retrieval task, not open-ended text generation.
Context contributed, but position was not enough
A contextual setup using context-bearing word targets and a sequence transformer showed a 19.8-point overall gain. The report’s decomposition put 5.9 points in context tracking and 13.9 points in context-independent gain.
A separate position probe found that position alone contributed 2.7 points of a 7.8-point within-run gain. After correcting for position, 5.1 points remained, including 3.1 points from the current trial and 1.9 points from preceding EEG.
The contextual setup also produced positive gains across rare, mid-frequency and frequent word groups: 6.0±3.5, 6.0±3.3 and 8.1±3.3 percentage points, respectively. The report notes that these gains may reflect decoded discourse context as well as direct information about the current word.
More data helped, while timing still mattered
The scaling analysis showed a log-linear relationship with training-data volume in the tested range. The contextual model gained 8.7 percentage points per decade of data, compared with 4.8 points for the non-contextual model; the fits had R² values of 0.98 and 0.99, and neither curve showed saturation.
Removing four occipital and posterior-temporal channels cut marginal within-run gain from 9.2±2.7 to 6.3±2.3 points, a relative drop of 32%. Context tracking fell by 4%. Because EEG activity projects across electrodes, the ablation does not cleanly locate the source of the signal.
Timing mattered as well. The gain was 6.9 points with a 0.5-second window, 8.3 with 0.7 seconds and 8.1 with 0.9 seconds; adding temporal jitter during training reduced performance by 13% at ±50 milliseconds and 39% at ±100 milliseconds. The pattern is compatible with time-locked activity, although it could also reflect limits in handling timing shifts.
Why the result remains preliminary
The authors interpret the findings as evidence that EEG carries word-level information during silent reading, including information related to the current trial and preceding EEG. But the study used one participant, and validation-set selection guided checkpoint and configuration choices, making the absolute gains upper bounds.
Silent reading in this one-word-at-a-time setup is only a proxy for inner speech; the result does not show that spontaneous inner speech can be decoded. The report says a 60-participant cross-subject cohort was still in preparation, which would test whether the finding generalizes beyond this participant.
Paper data and sources
Original title: Decoding silent reading from non-invasive EEG
Authors: Ingo Marquardt, Anthilia Alchanat, Priyanka Jain
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-20
DOI: Not available
Original paper · Full text