Toolkit for acoustic–phonetic analysis of naturalistic speech data
A major limitation in the speech sciences is access to naturalistic data in experimental settings and the difficulty of translating laboratory designs to real-world contexts. Researchers studying speech production or perception often rely on in-lab recordings, which constrain the sociolinguistic contexts examined and limit ecological validity. The Toolkit for Acoustic–Phonetic Analysis (TAPA) is an open-source pipeline that automates the acquisition, transcription, speaker diarization, forced alignment, and per-segment acoustic analysis of naturalistic, single/multi-speaker audio. The current release supports vowel formant extraction, stop voice onset time, and fricative spectral moments. We demonstrate TAPA on the 2016 U.S. presidential debate, extracting nearly 33,000 segments from a 90-min recording, and validate each measurement type against hand-coded annotation. Vowel formant agreement with expert measurements was high (F1 r = 0.89, F2 r = 0.87). Stop VOT showed reliable aggregate means but poor per-token agreement (r = − 0.04) because of a training–deployment mismatch in the neural VOT classifier. Fricative spectral standard deviation agreed strongly with hand-coded values overall (r = 0.80), and center of gravity agreed strongly for sibilants (/s/ r = 0.87, /ʃ/ r = 0.94), while non-sibilant moments diverged systematically. These findings suggest that TAPA can be used to increase access to naturalistic speech data and speed up the processing timeline with experts’ supervision.