SpeechLLMs have shown strong potential for emotion recognition, yet they read the predicted emotion off a generative decoder not suited for classification: it can emit labels outside the target set and favors frequent classes. We propose a discriminative adaptation that reads the final prompt token's hidden state throu...
These findings suggest that turn-taking is primarily conveyed through intonation and silence patterns rather than semantic completeness, enabling faster and more reliable systems without expensive text inference.
Rini A. Sharon, A. Manickavela, Kadri Hacioglu et al.· 0 citations
HOTFIXR: Hardness Optimized Training data For Improving X-Lingual Reasoning is a data generation framework that uses models to probe and learn a student model's multilingual weaknesses, and generates data to mitigate them that can improve multilingual performance.
Ishika Agarwal, Arkajyoti Charaborty, Tanner Sorensen et al.· 0 citations
This work shows that Group Relative Policy Optimization (GRPO) extracts far more from the same synthetic speech than SFT, and traces the gain to behavior rather than representation: GRPO reduces insertion errors by improving stopping calibration and speech-to-text alignment by better anchoring attention to audio, leavi...
Shashi Kumar, Yanis Labrak, Hasindri Watawana et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.