Skip to content
Open access

Visual speech enhances phoneme separability in human superior temporal gyrus

Sep 2026 · bioRxiv · 0 citations · 29 references
Biology Medicine

TL;DR

Using intracranial recordings from human superior temporal gyrus, it is shown that congruent visual speech does not simply amplify all speech-related information, instead, visual input selectively enhances the separability of confusable phoneme-level representations and improves word identity decoding, while leaving the robust phonetic-feature organization largely intact.

Abstract

Visual speech, such as lipreading, facilitates spoken word recognition, but the neural mechanisms underlying audiovisual speech perception remain poorly understood. Visual cues may disambiguate fine-grained articulatory features during early perceptual stages or instead integrate with speech at more categorical, phoneme-level stages. To test how speech representations are modulated by visual input, we analyzed intracranial electroencephalography (iEEG) signals recorded from 12 epilepsy patients performing an audiovisual speech perception task. Participants perceived 16 monosyllabic words presented in auditory-only, visual-only, or congruent audiovisual formats. Words were constructed from four onset consonants (/b/, /g/, /m/, /n/) and four rimes (vowel nucleus and any coda consonants). We examined event-related potentials (ERP) in superior temporal gyrus (STG) and trained support vector machine (SVM) classifiers to decode word identity from neural activity at individual electrodes. Discrete and continuous confusion matrices captured complementary changes in classification accuracy and normalized inverse classification loss, a continuous proxy for classifier confidence. Decoding performance was hierarchically evaluated at the word, phoneme, and phonetic feature levels to determine the representations affected by visual speech. Congruent audiovisual speech increased classifier confidence for phoneme-level representations and improved decoding accuracy at both the word and phoneme level, without corresponding effects on phonetic features. Time-resolved analyses further revealed earlier successful decoding for audiovisual than auditory-only speech, with audiovisual enhancement primarily observed for onset consonants rather than rimes. Together, these findings suggest that visual speech sharpens primarily categorical phoneme representations in STG, with accelerated speech processing and improved word recognition emerging as downstream consequences of phoneme-level enhancement. Significance Statement Seeing a speaker’s face improves speech perception, especially in noise, but the neural representations altered by visual speech remain unclear. Using intracranial recordings from human superior temporal gyrus, we show that congruent visual speech does not simply amplify all speech-related information. Instead, visual input selectively enhances the separability of confusable phoneme-level representations and improves word identity decoding, while leaving the robust phonetic-feature organization largely intact. These findings clarify the distinction between what visual speech modulates and the representational structure of speech in auditory cortex, suggesting that audiovisual facilitation primarily acts on categorical speech representations that more directly support word recognition.

Read PDF

Similar papers

Open access Sep 2026

Structural connectivity of auditory-linguistic brain networks predicts success in speech categorization and listening in noise

Successful speech perception requires listeners to bin continuous acoustic information into discrete phonetic categories. However, some people maintain within-category acoustic information (gradient) while others discard category-irrelevant information (discrete) during perception. Listeners also vary in how consistent...

Rose Rizzi, Jack R. Stirn, Zara Eisenhut et al. · 0 citations
Open access Aug 2026

Functional Distinctions between the Auditory Cortex and Frontoparietal Network in the Encoding of Speech Features

Neural representations of acoustic features that are key to speech perception have been studied extensively within the auditory cortex (AC), but their representations in higher-order cortical regions remain poorly understood. This study investigates the cortical encoding of three acoustic features in spoken vowels to d...

Yong-Tian Ou, K. Kay, A. Oxenham · 0 citations
Open access Sep 2026

Informational masking modulates semantic processing during recognition of syntactically complex speech: evidence from pupillometry

Introduction Speech perception in adverse listening environments depends not only on the quality of the acoustic signal but also on listeners’ ability to recruit higher-level linguistic information. Although semantic constraints are known to influence recognition of syntactically complex speech, it remains unclear how...

Jing Shen, G. DeDe · 0 citations
Preprint Aug 2026

Structured Phonological Representations for Audio-Articulatory rtMRI Speech Classification

Preliminary analyses show interpretable surface-sensitive patterns consistent with flapping-like /t/ realizations, /t/-/r/ retraction or affrication, and nasal place assimilation, indicating that phonological information from synchronized audio can be partially transferred to articulatory models.

Abner Hernandez, T. A. Vergara, Dai-Qi Liu et al. · 0 citations

Thai speech envelope auditory attention detection using eeg

Auditory attention detection (AAD) identifies which of several competing talkers a listener is attending to, a key step toward neuro-steered hearing devices for real-world listening environments with multiple speakers. Most AAD work to date has examined non-tonal languages, leaving tonal languages underexplored even th...

Shalong Samretngan · 0 citations
Open access Sep 2026

Voice and linguistic features support early attentional filtering of irrelevant stimuli

In cocktail party-type environments, listeners must segregate simultaneous acoustic sources and select one for processing. Behavioral studies have established that both low-level voice differences (e.g., pitch) and high-level linguistic differences (e.g., word content) aid these processes. Electroencephalography (EEG)...

Sahil Luthra, Eric Parker, Marysia Brown et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.