Preprint
Aug 2026
CASA: Content-Acoustic Speaking Assessment with Speech Encoder and Large Language Model
CASA, a simpler architecture combining Whisper-medium and Qwen3.5-2B that achieves state-of-the-art performance while providing a more interpretable separation between speech delivery and content, is proposed.
Nhan Phan, Ilona Lähteenmäki, Anna von Zansen et al.
· 0 citations