NSV-Shift: A Contrastive Benchmark for Non-Speech Vocalization Understanding and Response Adaptation in Speech-to-Speech Models
This work introduces NSV-Shift, a contrastive benchmark for evaluating whether speech-to-speech models can understand non-speech vocalizations (NSVs) and adapt their responses accordingly, and evaluates five models on NSV perception, emotion understanding, and response adaptation.
Zi-Wei Chen
· 0 citations