Skip to content
Open access

Voice and linguistic features support early attentional filtering of irrelevant stimuli

Sep 2026 · bioRxiv · 0 citations · 31 references
Medicine Biology

Abstract

In cocktail party-type environments, listeners must segregate simultaneous acoustic sources and select one for processing. Behavioral studies have established that both low-level voice differences (e.g., pitch) and high-level linguistic differences (e.g., word content) aid these processes. Electroencephalography (EEG) studies have demonstrated that the brain filters voices based on low-level features early in sensory processing, but it is unclear whether linguistic differences also support early filtering. In this EEG study, listeners attended to a lateralized auditory stream of target syllables while ignoring distractor stimuli from the opposite hemifield. In contrast to previous auditory attention studies, target stimuli were always attended and statistically identical across conditions; instead, we manipulated the distractors. Specifically, we manipulated whether targets and distractors were spoken in the same voice or different voices, as well as whether they were composed of the same or different linguistic content. Critically, distractors were temporally offset from targets, enabling an investigation of whether acoustically matched, always attended target stimuli are differentially encoded as a function of surrounding context. Differences in either voice or linguistic features supported target recall: Recall was poor only when both streams were in the same voice and had similar linguistic content. However, voice and linguistic features had independent, additive effects on neural processing. Specifically, target dissimilarity in acoustics and linguistic content both contributed to more robust, larger-magnitude target-evoked EEG responses. Overall, results demonstrate that both linguistic dissimilarity and low-level acoustic dissimilarity improve early attentional filtering of distractors in multi-source acoustic scenes. Graphical Abstract Highlights We varied the voice and semantic content of distractors in a cocktail party task Auditory targets were statistically identical and always attended across conditions Target recall suffered only when streams were similar in both voice and content But voice and linguistic features independently shaped attentional EEG responses Similarity in either voice or content weakened target-evoked neural responses

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.