Skip to content
Open access

Inner speech decoding from EEG

Abstract

Inner speech the silent production of words in the mind, without any movement or sound is an appealing control signal for a brain computer interface (BCI), because the command is the thought. For someone who has lost the ability to speak or move, a decoder that reads intended words directly would be far more natural than the indirect mental tasks most BCIs rely on. Reading inner speech from scalp electroencephalography (EEG) is, however, extremely hard: the signals are weak and non-stationary, the neural traces of covert language are faint, and reported four-class accuracies in the literature rarely climb much past the low thirties. In this regime the meaningful question is not whether a model reaches high accuracy none reliably do - but whether a design choice yields a real, statistically reliable signal, assessed honestly. This thesis compares three standard architectures a compact convolutional network (EEGNet), an LSTM recurrent network, and a self-attention Transformer against a hybrid model that feeds a shared convolutional front end into parallel recurrent and self-attention branches and fuses them before classification. All four are evaluated identically on the public Thinking Out Loud inner speech dataset (Nieto et al., 2022), under subject-dependent five-fold cross-validation on the four-class directional-word task (Up, Down, Right, Left), with on-the-fly augmentation and a fixed seed. The unit of statistical analysis is the subject (n = 10). The hybrid model attains the highest mean accuracy, 28.8% (SD 3.0%), and is the only model whose accuracy is statistically significantly above the 25% chance level (Wilcoxon signed-rank p = 0.010; one-sample t-test p = 0.003), a result that survives Bonferroni correction across the four model-versus-chance tests. The three baselines do not reach significance against chance. In direct paired comparisons, however, the hybrid does not significantly outperform any individual baseline (versus EEGNet p = 0.064; versus the Transformer p = 0.084; versus the LSTM p = 0.125), and an ablation shows that each single branch performs at roughly the level of its corresponding baseline (≈ 0.26), with the two branch fusion adding a small, non-significant improvement. With ten subjects, the study can resolve only differences of about four percentage points or larger, so these null results are inconclusive rather than evidence of equivalence. The protocol was executed twice on identical hardware: the baselines reproduced exactly, while the hybrid’s mean shifted by two tenths of a percentage point enough to move its comparison against EEGNet from p = 0.014 to p = 0.064. The above-chance result held in both executions; the paired comparisons did not, and are reported here as unstable at this sample size. Only EEGNet was evaluated under a subject-independent split, where it scored 25.6-26.7%; because that baseline does not exceed chance under subject-dependent evaluation either, these experiments do not measure the size of the cross-subject gap, which is cited here from the literature rather than demonstrated. The accuracies reported here fall within the range of prior decoding studies on this dataset, which spans roughly 26-36%. The contribution of this thesis is therefore a controlled, like-for-like benchmark of four architectures under one protocol, with rigorous statistical assessment, rather than a new state of the art.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.