Through a four-speaker attention decoding benchmark, it is shown that combining behavioral and physiological signals improves decoding performance over EEG-only approaches, enabling future advances in multimodal auditory attention decoding.
Abstract
Humans rely on gaze, head movements, and visual cues to attend to speakers in noisy environments, yet auditory attention decoding (AAD) has been studied primarily using electroencephalography (EEG). We introduce the Multimodal Auditory-attention Egocentric Speech-TRacking Open (MAESTRO) corpus, the first AAD dataset to simultaneously record EEG, eye gaze, pupillometry, egocentric video, and head inertial measurement unit (IMU) data. MAESTRO includes four competing speakers and background noise across multiple signal-to-noise ratio (SNR) conditions, enabling attention decoding under realistic listening scenarios. Through a four-speaker attention decoding benchmark, we show that combining behavioral and physiological signals improves decoding performance over EEG-only approaches, enabling future advances in multimodal auditory attention decoding. These findings open the door to new applications, analyses, and methodological advances in multimodal AAD. The complete dataset is publicly available at https://huggingface.co/datasets/aspire-osu/maestro-eeg-dataset . The official code repository is available at https://github.com/ASPIRE-OSU/MAESTRO .
Auditory attention decoding (AAD) is often evaluated on static, simplified speech scenes that poorly match everyday listening. We introduce MOV-AAD, a large-scale dataset for studying auditory attention under moving, naturalistic conversations. MOV-AAD combines 64-channel EEG with synchronized physiological recordings,...
Xiao-Min He, Vishal Choudhari, Tristan J. Spratt et al.· 0 citations
Auditory attention detection (AAD) identifies which of several competing talkers a listener is attending to, a key step toward neuro-steered hearing devices for real-world listening environments with multiple speakers. Most AAD work to date has examined non-tonal languages, leaving tonal languages underexplored even th...
A dual-stream time-frequency convolutional network with additive attention that improves sensitivity to and modeling of frequency variation patterns and significantly outperform state-of-the-art methods is proposed.
Xiao Lin, Sunjie Zhang, Shuai Huang· Sheng wu yi xue gong cheng x...· 0 citations
Objective electrophysiological markers are often proposed for speech-in-noise assessment, but their validation targets are not always explicit: a marker may calibrate with intelligibility, distinguish acoustic conditions, decode attention, or predict an individual speech reception threshold (SRT). We used three public...
Li Guo, Yang Li, Jun-Lin Wang et al.· Hearing Research· 0 citations
Stimulus-driven auditory attention determines which sound gains priority when multiple sources compete without an explicit listening goal, yet most computational studies focus either on acoustic salience or on decoding predefined attended targets. This study investigates instruction-free auditory competition using the...
During audiovisual perception, spatial information from vision and audition is combined, often producing biases such as the ventriloquist effect. While these interactions are well documented behaviourally, it remains unclear when cross-modal information begins to alter modality-specific spatial representations in the b...
Zak Buhmann, Amanda K. Robinson, Jason B. Mattingley et al.· Journal of Neuroscience· 0 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduOct 7, 2026
Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.
MIT News · Artificial Intelligence· news.mit.eduSep 25, 2026