Multimodal Large Language Models (MLLMs) provide a natural way to make video anomaly detection more explainable. However, their final decisions do not always fully use the discriminative information contained in their hidden states, an issue we refer to as representation--behavior misalignment. We decompose this gap in...
Chao Huang, Peng-Fei Wei, Kai-Ge Li et al.· 0 citations
RoMod is proposed, an efficient VAD framework trained with only \(5\%\) of weakly labeled videos that achieves state-of-the-art performance while running substantially faster than dense backbones of comparable size.
Chao Huang, Peng-Fei Wei, Ben-Feng Wang et al.· 0 citations
Accurate medical image segmentation plays a vital role in clinical diagnostics by facilitating the precise delineation of anatomical structures and pathological regions. However, the performance of existing segmentation methods is often constrained by the scarcity of high-quality annotated datasets, as manual labeling...
Chao Huang, Peng Chen, Jie Wen et al.· IEEE Transactions on Image P...· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.