What Did the MLLM Hear? Token-Level Spectro-Temporal Grounding for Audio MLLM Explainability
STAG is introduced, to the authors' knowledge the first post-hoc framework for token-level spectro-temporal grounding of captions generated by audio-based MLLMs, and provides behavioral support for the faithfulness and selectivity of the explanations.
Lucia Cascone, V. Fraenza, Michele Nappi et al.
· 0 citations