SpanCalib-VLM: Calibrated Hallucination Span Detection in Vision-Language Models
SpanCalib-VLM is presented, a hybrid dual-system for the SHROOM-Visions Shared Task that combines a multimodal sequence tagger, consisting of XLM-RoBERTa-Large fused with a SigLIP vision encoder via cross-attention, with the fine-tuned generative VLM (Qwen3.5-4B-SHROOM-SFT).
A. Abebe, Yasmin Moslem
· 0 citations