Video multimodal large language models (MLLMs) keep climbing video question answering benchmarks, yet shuffling the frames, masking the segment that supports the answer, or occluding the target object barely changes their predictions. The accuracy rests on appearance and language priors, not on the temporal evidence th...
Zhaolu Kang, Shi-Yu Liu, Tai-Long Luo et al.· 0 citations
Savor is introduced, a training framework that augments the output schema with token and answer confidence, optimises the policy with a Group Relative Policy Optimisation objective that penalises calibration error and poor abstention decisions, and uses the learned confidence at inference time to revisit visual evidenc...
Zian Ding, Zi-Lin Zhao, Ying-Jie He et al.· 0 citations
OrderProbe is introduced, a deterministic benchmark for structural reconstruction using fixed four-character expressions in Chinese, Japanese, and Korean, which have a unique canonical order and thus support exact-match scoring.