Auditing Generative Audio Calls for Known-Task Audio-LLM Evaluation
This work evaluates speech and audio LLMs as a controlled call-decision problem, and finds that agreement and stacking features improve weaker selectors but do not beat the strongest no-call control.