Skip to content

Author

Caifeng Shan

We have 3 of 6 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#small language model Preprint Sep 2026

Probe-VAD: Ordinal Likelihood Probing for Training-Free Video Anomaly Detection

Probe-VAD is proposed, an ordinal binary-probing framework that directly probes severity preferences from a frozen VLM, providing a simple interface for translating frozen VLM visual understanding into continuous, rank-sensitive anomaly scores without task-specific training or caption-based compression.

Jia-Wei Gu, Qi-Lin Zhao, Teng-Kuo Guo et al. · 0 citations
Preprint Jul 2026

Global Logic and Local Search: Dual-Stream Multimodal In-Context Learning for Verifiable Industrial Anomaly Detection

Large Multimodal Models (LMMs) show strong few-shot generalization, but industrial anomaly detection remains difficult because defects are small, input resolution is limited, and textual standards are not always grounded in visual evidence. Recent optimization-based methods improve alignment through fine-tuning, but they often require many defective samples, which are unavailable in early deployment. We present Global Logic and Local Search (GLLS), a training-free framework for reference-guided multimodal in-context verification. GLLS uses a Part-Aware Visual-Logical Atlas to organize normal references and structured specifications in the inference context. It combines a Global&Logic Stream, where SAM 3 extracts partially checkable visual facts, with a Fine-Grained&Actions Stream, where MCTS selects local evidence crops under a fixed budget. Experiments on MMAD-QA and additional anomaly detection datasets show consistent gains over matched and general-purpose baselines, while keeping the final diagnostic decision traceable to explicit visual evidence throughout the inspection trace.

R. Deng, Yundi Hu, Yiming Zhong et al. · 0 citations
Preprint Aug 2026

TD-VAD: Breaking Visual Dependence in Video Anomaly Detection with Text-Driven Learning

This work proposes a novel Text-Driven Video Anomaly Detection (TD-VAD) approach, which utilizes video-like text descriptions with temporal characteristics generated by LLM to train a VAD model, without any reliance on target-domain anomaly data.

Shuang-Qing Zhang, Lei-Lei Ma, Zhao Wang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.