It is demonstrated that an end-to-end deep learning framework can yield promising decoding performance from dynamic visual stimuli without handcrafted features, while the model behavior remains interpretable through spectral, temporal, and cortical dimensions, which are broadly consistent with established neuroscience knowledge.
Abstract
ECoG-based visual semantic decoding enables inference of semantic interpretation of visual perception from complex, noisy brain activity. This study examines the feasibility of visual semantic decoding using an end-to-end deep learning framework using electrocorticography (ECoG). Specifically, the decoding task is to predict visual categories from video stimuli using time-series neural inputs. A previously collected ECoG dataset from participants ($n=17$) with drug-resistant epilepsy is used for analysis. With fewer than 50 training samples per visual category, this study evaluates multiple deep learning approaches, artificial neural network architectures, and frequency-band filtered inputs. The best-performing approach is analyzed to shed light on the discriminative information it relies on across spectral, temporal, and cortical dimensions. The selected decoding system uses mixup augmentation, a Transformer-based encoder, and high-gamma (80-150 Hz) inputs with a 900 ms post-stimulus window. Further analysis shows that early visual cortex (V2-V4), ventral stream visual cortex, MT+ complex with neighbouring visual areas, and lateral temporal cortex contributed substantially to decoding performance. This study demonstrates that an end-to-end deep learning framework can yield promising decoding performance from dynamic visual stimuli without handcrafted features, while the model behavior remains interpretable through spectral, temporal, and cortical dimensions, which are broadly consistent with established neuroscience knowledge.
Paired MEG occlusion shows that 15 of 19 stimulus features contribute, with the largest effects for silence, sound intensity, vowels, and acoustic onsets, indicating that activity without narrative structure carries less recoverable information than activity during coherent speech.
Ilia Semenkov, Daria Kleeva, I. Dakhtin et al.· 0 citations
FuzzyAlign, an alignment framework driven by fuzzy similarity, is proposed to establish a benchmark and explore the integration of large pretrained vision models with neural decoding, offering a high-performing and interpretable approach for bridging neural and artificial vision systems.
Yonghao Song, Chengjian Xu, Qingqing Zheng et al.· IEEE transactions on fuzzy s...· 0 citations
By aligning EEG with layer-wise neural visibility rather than fixed high-level semantics, the proposed framework improves both retrieval accuracy and image reconstruction in EEG-based visual decoding.
Decoding visual information based on machine learning has the potential to reveal the neural mechanisms of visual information processing in the brain. Although most highly precise decoding methods require highly invasive electrodes, these electrodes make long-term use difficult because they cause brain damage. To addre...
A comprehensive taxonomy of MI EEG cross-variability decoding studies from 2020 to 2025 is presented, systematically organizing advances in deep learning and transfer learning and critically evaluate core algorithmic approaches, including Convolutional Neural Networks, transformers, feature alignment, domain adaptation...
Li-Jun Wang, Yue-Ying Zhou, Peng-Pai Wang et al.· Frontiers in Neuroscience· 0 citations
Electroencephalography-based visual decoding has important applications in brain–computer interfaces and cognitive neuroscience, yet the relative effectiveness of different feature extraction methods for sustained visual paradigms remains unclear due to the absence of standardized, multi-dataset comparative evaluations...
Cesar-Agustin Corona-Patricio, Carolina Reta, J. A. Cantoral-Ceballos· Applied Informatics· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.