Skip to content

Accurate Decoding of Natural Sentences from Non-Invasive Brain Recordings

Jun 2026 · 0 citations · 58 references
Computer Science Engineering Biology

TL;DR

The results show that non-invasive brain-to-text decoding starts to operate at a level of accuracy previously thought exclusive to surgical implants, opening a path toward safe and efficient brain-computer-interfaces.

Abstract

Restoring communication for people who have lost the ability to speak or move after a brain injury is a major challenge. While intracranial implants now enable high-performing brain-computer-interfaces, non-invasive alternatives are still lagging behind. Here, we present Brain2Qwerty v2, a model that can decode the production of natural sentences solely from real-time magnetoencephalography (MEG) recordings. By collecting 22,000 sentences typed by nine subjects, each recorded for 10 hours, our model leverages character, word and sentence-level representations to achieve an average word error rate (WER) of 39%. For our best participant, the model accurately decodes half of the sentences with one word error or less. Critically, decoding accuracy log-linearly improves with data volume, suggesting that the performance gap with intracranial approaches could be partially bridged through data scaling. We show that AI enables this performance in three main ways: the substitution of hand-crafted pipelines for event detection with deep learning, the finetuning of large language models to extract semantic representations, and the deployment of AI agents to iteratively refine our decoding pipeline via automated code development. Together, these results show that non-invasive brain-to-text decoding starts to operate at a level of accuracy previously thought exclusive to surgical implants, opening a path toward safe and efficient brain-computer-interfaces.

View source

Similar papers

Preprint Aug 2026

Decoding silent reading from non-invasive EEG

Non-invasive decoding of inner speech faces a fundamental data problem: a corpus pairing brain activity with a person's spontaneous inner monologue cannot be collected, and the available proxy paradigms (cued repetitive and retrospectively reported generative inner speech) are slow to acquire, poorly time-locked, and subject compliance is unverifiable. We therefore treat silent reading as a scalable proxy task and ask how much lexical and semantic information a contrastive decoder can extract from it. We report an open-vocabulary analysis of approximately 240,000 word presentations recorded from a single densely-sampled participant across 393 runs (ca. 49 h) of 19-channel dry-electrode EEG. Words from continuous narrative text were presented in rapid serial visual presentation, with typography randomised on every trial to partially decorrelate word identity from low-level visual form. A convolutional EEG encoder, optionally followed by a causal transformer, was trained with a CLIP-style contrastive objective to align short EEG windows with hidden-state embeddings of the presented word taken from a large language model. Decoding, evaluated as word-grouped top-10 retrieval against permutation baselines, was reliably above chance, extended to mid-frequency and rare words, and scaled log-linearly with training-data volume with no sign of saturation. Removing occipital and posterior-temporal electrodes reduced the word-level gain by roughly one third but left context tracking unchanged. Control analyses separate word-level decoding from narrative context tracking and from a non-neural positional prior introduced by the transformer's positional embedding. These results establish that open-vocabulary word-level information is recoverable from EEG during silent reading, and that decoding is data-limited rather than saturated.

I. Marquardt, A. Alchanat, Priyanka Jain · 0 citations
Open access Jul 2026

A generalizable speech neuroprosthesis

Zachery M. Fogg, N. Card, M. Wairagkar et al. · 0 citations
Preprint Aug 2026

Interpretable MEG Decoding of Perceived Speech: Cortical Sources and the Stimulus Features That Drive Retrieval

Short segments of perceived speech can be retrieved from non-invasive magnetoencephalographic (MEG) recordings by deep networks trained with a CLIP-style objective against wav2vec 2.0 audio embeddings. Yet their weights do not map onto electrophysiological quantities, and it remains unclear which speech properties drive retrieval. We build on a high-performing MEG-to-audio retrieval architecture but redesign both its front end and decoder. Its spatial attention operates on a flattened sensor layout; we replace it with spherical harmonics defined on the three-dimensional MEG helmet geometry. We reduce the subject-specific representation from 270 to 25 branches, add a temporal filter to each branch to match it to a neuronal source in space and time, and make the convolutional decoder shallower. Ocular and cardiac components are removed before training to reduce the risk of stimulus-locked shortcuts. On MEG-MASC, the model reaches 39.75 +/- 0.34% Top-1 accuracy among 1005 candidates across six trained solutions, with about 20 times fewer decoder parameters. Its weights map to source space, recovering generators consistent with the speech-perception network, while left-lateralized branches carry higher-frequency rhythmic components not evident on the right. Paired MEG occlusion shows that 15 of 19 stimulus features contribute, with the largest effects for silence, sound intensity, vowels, and acoustic onsets. Random word lists behave oppositely: substituting narrative MEG into them improves retrieval, indicating that activity without narrative structure carries less recoverable information than activity during coherent speech. The wav2vec target can be reduced to about twelve learned feature dimensions without loss of accuracy, whereas strong temporal compression causes a clear loss. Together, source mapping and input interventions reveal what drives retrieval.

Ilia Semenkov, Daria Kleeva, I. Dakhtin et al. · 0 citations
Jul 2026

Natural language decoding from EEG via transfer learning and multimodal contrastive learning

Decoding natural language from non-invasive electroencephalography (EEG) is a key step toward practical brain–computer interfaces for individuals with speech impairments. While prior work on the Large Spanish Speech EEG Dataset has focused on sentence classification, semantic reconstruction remains largely unexplored. We propose a framework for sentence-level language decoding based on transfer learning and multimodal contrastive alignment, training an EEG encoder to align neural signals with pretrained text embeddings that are subsequently inverted into text using a pretrained embedding-to-text model. Under 10-fold subject-wise cross-validation, our approach achieves a BERTScore of 0.320 ± 0.021, BLEU-4 of 0.249 ± 0.026, and 21.7 ± 0.031% sentence classification accuracy over 30 classes. Ablation results indicate that reconstruction performance is primarily driven by contrastive alignment. These findings demonstrate the feasibility of semantically meaningful sentence reconstruction from non-invasive EEG.

Jose Manuel Carrichi Chavez, Toru Nakashika, Tomoaki Mizuno et al. · 0 citations
Open access Jul 2026

Shared latent representations of speech production for cross-patient speech decoding

Speech brain-computer interfaces (BCIs) can restore communication in individuals with neuromotor disorders who are unable to speak. However, current speech BCIs limit patient usability and successful deployment by requiring large volumes of patient-specific data collected over long periods of time. A promising solution to facilitate usability and accelerate their successful deployment is to combine data from multiple patients. This has proven difficult, however, due to differences in user neuroanatomy, varied placement of electrode arrays, and sparse sampling of targeted anatomy. Here, by aligning patient-specific neural data to a shared latent space, we show that speech BCIs can be trained on data combined across patients. Using canonical correlation analysis and high-density micro-electrocorticography (μECoG), we uncovered shared neural latent dynamics with preserved micro-scale speech information. This approach enabled cross-patient decoding models to achieve improved performance relative to patient-specific models facilitated by the high resolution and broad coverage of μECoG. Our findings support future speech BCIs that are more accurate and rapidly deployable, ultimately improving the quality of life for people with impaired communication from neuromotor disorders. Current speech brain-computer interfaces (BCIs) rely on patient-specific decoding approaches. Here, the authors show that patient-specific data can be aligned to a shared space that preserves speech information, enabling cross-patient speech BCIs.

Z. Spalding, S. Duraivel, S. Rahimpour et al. · 0 citations
Preprint Aug 2026

LibriBrain100: One Hundred Hours of Broad and Deep MEG Data for Neural Speech Decoding at Scale

We introduce LibriBrain100, a large-scale MEG dataset for speech decoding designed from the ground up for reproducible, standardised evaluation. LibriBrain100 more than doubles the size of the original LibriBrain release, resulting in over 100 hours of high-quality MEG acquired while subjects listened to naturalistic continuous speech. With $\sim$80 hours from a single subject, LibriBrain100 sets a new record for deep, within-subject neural data (8$\times$ more than the next comparable dataset and roughly 80$\times$ more than other datasets). To demonstrate the payoff of this depth-first design, we evaluate on a word-classification benchmark---an increasingly well-established stepping stone towards the open challenge of noninvasive brain-to-text decoding. Using an existing decoding model, we achieve state-of-the-art performance---validating both the quality of the recordings and the value of within-subject data at scale. Because collecting 80 hours of data per user is impractical for real-world applications, we also collected $\sim$40 minutes of additional data from each of 32 subjects. Using the same word-classification benchmark, we demonstrate the value of broad multi-subject data: supervised finetuning of a pre-trained model can substantially compensate for limited per-subject data. We provide standard train, validation, and test splits, all reproducible through an open-sourced Python library that supports easy downloading, optional preprocessing, and data loading for common deep learning frameworks. In addition, the dataset and evaluation infrastructure are being released alongside an open machine-learning competition with a public leaderboard for standardised benchmarking. Ultimately, our hope is that LibriBrain100 will accelerate progress towards practical non-invasive brain-computer interfaces, capable of restoring communication to people living with severe paralysis.

Francesco Mantegna, D. Jayalath, Gereon Elvers et al. · 0 citations

Related blog posts