Thai speech envelope auditory attention detection using eeg
Abstract
Auditory attention detection (AAD) identifies which of several competing talkers a listener is attending to, a key step toward neuro-steered hearing devices for real-world listening environments with multiple speakers. Most AAD work to date has examined non-tonal languages, leaving tonal languages underexplored even though they account for over half of the world's languages. This study addresses that gap using Thai, a tonal language with five lexical tones. Electroencephalography (EEG) was acquired from 30 normal-hearing native speakers through a wearable 8-channel montage. The subjects attended to one of two competing Thai podcast streams, played simultaneously through two-channel over-ear headphones with one stream to each ear, throughout each trial. Two decoders were compared: a linear stimulus reconstruction (LSR) model combining EEG and speech, and an EEG-only convolutional neural network (CNN) evaluated with Dropout and SpatialDropout regularization. The LSR decoder achieved above-chance average accuracy at every decision window from 1 to 60 s (50.24 to 51.39% for all 30 subjects; 50.79 to 55.40% for the 21 LSR-decodable subjects). Four speech-envelope extraction methods, designed to better represent Thai's tonal frequencies, showed no significant differences in LSR performance. The CNN, which requires more training data and is therefore weaker at short decision windows, exceeded LSR only at the 60 s window (CNN-Dropout 56.44% versus LSR 55.40% in the LSR-decodable subgroup). The study confirms that tonal-language AAD is feasible with a low-channel wearable setup and releases an open Thai AAD EEG dataset.