Using intracranial recordings from human superior temporal gyrus, it is shown that congruent visual speech does not simply amplify all speech-related information, instead, visual input selectively enhances the separability of confusable phoneme-level representations and improves word identity decoding, while leaving the robust phonetic-feature organization largely intact.
Abstract
Visual speech, such as lipreading, facilitates spoken word recognition, but the neural mechanisms underlying audiovisual speech perception remain poorly understood. Visual cues may disambiguate fine-grained articulatory features during early perceptual stages or instead integrate with speech at more categorical, phoneme-level stages. To test how speech representations are modulated by visual input, we analyzed intracranial electroencephalography (iEEG) signals recorded from 12 epilepsy patients performing an audiovisual speech perception task. Participants perceived 16 monosyllabic words presented in auditory-only, visual-only, or congruent audiovisual formats. Words were constructed from four onset consonants (/b/, /g/, /m/, /n/) and four rimes (vowel nucleus and any coda consonants). We examined event-related potentials (ERP) in superior temporal gyrus (STG) and trained support vector machine (SVM) classifiers to decode word identity from neural activity at individual electrodes. Discrete and continuous confusion matrices captured complementary changes in classification accuracy and normalized inverse classification loss, a continuous proxy for classifier confidence. Decoding performance was hierarchically evaluated at the word, phoneme, and phonetic feature levels to determine the representations affected by visual speech. Congruent audiovisual speech increased classifier confidence for phoneme-level representations and improved decoding accuracy at both the word and phoneme level, without corresponding effects on phonetic features. Time-resolved analyses further revealed earlier successful decoding for audiovisual than auditory-only speech, with audiovisual enhancement primarily observed for onset consonants rather than rimes. Together, these findings suggest that visual speech sharpens primarily categorical phoneme representations in STG, with accelerated speech processing and improved word recognition emerging as downstream consequences of phoneme-level enhancement. Significance Statement Seeing a speaker’s face improves speech perception, especially in noise, but the neural representations altered by visual speech remain unclear. Using intracranial recordings from human superior temporal gyrus, we show that congruent visual speech does not simply amplify all speech-related information. Instead, visual input selectively enhances the separability of confusable phoneme-level representations and improves word identity decoding, while leaving the robust phonetic-feature organization largely intact. These findings clarify the distinction between what visual speech modulates and the representational structure of speech in auditory cortex, suggesting that audiovisual facilitation primarily acts on categorical speech representations that more directly support word recognition.
Successful speech perception requires listeners to bin continuous acoustic information into discrete phonetic categories. However, some people maintain within-category acoustic information (gradient) while others discard category-irrelevant information (discrete) during perception. Listeners also vary in how consistent...
Rose Rizzi, Jack R. Stirn, Zara Eisenhut et al.· Neuroscience· 0 citations
Neural representations of acoustic features that are key to speech perception have been studied extensively within the auditory cortex (AC), but their representations in higher-order cortical regions remain poorly understood. This study investigates the cortical encoding of three acoustic features in spoken vowels to d...
Yong-Tian Ou, K. Kay, A. Oxenham· Journal of Neuroscience· 0 citations
Introduction Speech perception in adverse listening environments depends not only on the quality of the acoustic signal but also on listeners’ ability to recruit higher-level linguistic information. Although semantic constraints are known to influence recognition of syntactically complex speech, it remains unclear how...
Jing Shen, G. DeDe· Frontiers in Psychology· 0 citations
Preliminary analyses show interpretable surface-sensitive patterns consistent with flapping-like /t/ realizations, /t/-/r/ retraction or affrication, and nasal place assimilation, indicating that phonological information from synchronized audio can be partially transferred to articulatory models.
Abner Hernandez, T. A. Vergara, Dai-Qi Liu et al.· 0 citations
Auditory attention detection (AAD) identifies which of several competing talkers a listener is attending to, a key step toward neuro-steered hearing devices for real-world listening environments with multiple speakers. Most AAD work to date has examined non-tonal languages, leaving tonal languages underexplored even th...
In cocktail party-type environments, listeners must segregate simultaneous acoustic sources and select one for processing. Behavioral studies have established that both low-level voice differences (e.g., pitch) and high-level linguistic differences (e.g., word content) aid these processes. Electroencephalography (EEG)...
Sahil Luthra, Eric Parker, Marysia Brown et al.· bioRxiv· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.