Audio language models can recognize spoken commands and invoke tools, but an agent must first decide whether the acoustic and conversational context warrants action. We introduce VGBench, a 1,018-item diagnostic benchmark for action-level addressedness across side-talk, self-talk, and speaker-switch scenarios. Each ite...
Intent drift is established as a measurable multi-turn failure mode and explicit state maintenance as a partial mitigation and IntentFlux is introduced, an executable benchmark that converts verifiable tasks into dialogues with controlled intent changes while preserving their original graders.
Yan-Jie Zhang, Bo-Wen Cao, Zi-Xin Chen et al.· 0 citations
Memory-augmented VLM agents act on persistent spatial knowledge, yet that knowledge silently goes stale as the environment changes. We ask what happens when an agent must reconcile a confident memory claim with a contradicting observation, and whether current models can catch the conflict before it becomes a safety-rel...
Yu-Shi Sun, Yan-Jie Zhang· 1 citation· ⚡1
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.