Multimodal Large Language Models (MLLMs) provide a natural way to make video anomaly detection more explainable. However, their final decisions do not always fully use the discriminative information contained in their hidden states, an issue we refer to as representation--behavior misalignment. We decompose this gap in...
Chao Huang, Peng-Fei Wei, Kai-Ge Li et al.· 0 citations
Out-of-context (OOC) misinformation, where genuine images are paired with misleading text, poses substantial societal risks due to its deceptive narratives. Prior efforts primarily assess image-text consistency but often lack explainable judgments, limiting their usefulness for forensic validation. Recently, Multimodal...
Kai-Ge Li, Zhao-Wei Wu, Chao Huang et al.· IEEE Transactions on Image P...· 0 citations
This work introduces a Variational Semantic Prompt Extractor (VSPE), which adaptively aggregates anomaly-relevant local semantics from dense patch tokens and regularizes them through a variational information bottleneck, thereby incorporating fine-grained visual cues and enabling more precise cross-modal alignment.
Peng Chen, Kai-Ge Li, Wei Wang et al.· arXiv.org· 0 citations
AdvNav is proposed, a behavior-guided black-box adversarial attack framework that disturbs an agent's first-person views during navigation, which demonstrates the effectiveness and generality of AdvNav, reveals critical perception vulnerabilities and offers insights for the design of future resilient VLN models.
Chenyang Li, Kaige Li, Zeyu Jiang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.