In-context learning (ICL) promises training-free adaptation for audio, where labeling every new condition is costly. Yet existing audio ICL studies largely measure Task Recognition, where demonstrations merely cue pre-trained capabilities, rather than Task Learning, where a genuinely new input-label mapping must be inferred from demonstrations alone. We introduce AudioICL-Bench, a diagnostic benchmark whose per-episode rules are resampled so that no correct answer is recoverable from prior knowledge. Its nine tasks are organized along two axes that separate what must be learned from demonstrations from what must be perceived in the signal, enabling failures to be attributed to either source. Across five Large Audio Language Models, the strongest models readily bind arbitrary sounds to new labels when perception is easy, but collapse on temporal measurement and composing multiple induced rules, revealing two primary capability boundaries: temporal perception and multi-rule composition.
Jia-Hung Chen, Yi-Cheng Lin, Kai-Wei Chang et al.· 0 citations
The Teaching Monster Challenge is introduced, the first instructional video generation benchmark to treat the learner persona as an explicit evaluation criterion and release the benchmark, rubric, and human judgments as a testbed for both teaching systems and automatic judges.
Yi-Cheng Lin, Yu-Kai Guo, Szu-Chi Chen et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.