When Demonstrations Fail: Diagnosing the Limits of In-Context Learning in Large Audio-Language Models with Progressive Cue Removal
A three-stage evaluation pipeline is designed that progressively reduces textual guidance to systematically evaluate LALMs' in-context learning ability in the audio modality, uncovering a consistent asymmetry across LALMs.