Skip to content

IoUPD: IoU-Aware Privileged Distillation for Visual Grounding with Multimodal Large Language Models

Jul 2026 · arXiv.org · Vol abs/2607.15732 · 0 citations · 30 references
Computer Science

TL;DR

Experiments on standard referring-expression grounding benchmarks show consistent region-level improvements over strong coordinate-generating baselines, demonstrating that ground-truth boxes can provide useful privileged guidance beyond serving as coordinate labels.

Abstract

Visual grounding with multimodal large language models is commonly formulated as autoregressive coordinate generation, where a model outputs bounding-box coordinates as text given an image and a referring-expression prompt. While this interface is simple and compatible with instruction following, it introduces a mismatch between training and evaluation: training optimizes token-level likelihood over coordinate strings, whereas grounding quality is measured by geometric overlap. We propose IoU-PD, an IoU-aware privileged distillation method for coordinate-generating multimodal large language models. IoU-PD uses ground-truth boxes not only as coordinate targets, but also as privileged training-time guidance. During training, the student receives the original image and prompt, while a frozen teacher receives a box-marked image and an augmented prompt that indicates the marked region. The student is trained with a supervised fine-tuning anchor and a privileged distillation loss whose token weights reflect both geometric importance and teacher reliability. At inference time, IoU-PD requires no box overlay, privileged hint, teacher branch, or additional prediction module. Experiments on standard referring-expression grounding benchmarks show consistent region-level improvements over strong coordinate-generating baselines, demonstrating that ground-truth boxes can provide useful privileged guidance beyond serving as coordinate labels. Project page: https://xyzzzh.github.io/IoU-PD/

View source

Similar papers

Preprint Aug 2026

Hi-Token: Hierarchical Coordinate Tokenization for Generative Visual Grounding

The results show that structured coordinate generation provides an effective approach to generative visual grounding and Hi-Token encodes each coordinate with axis-specific tokens for the hundreds, tens, and ones digits, which adds coarse-to-fine structure and increases token reuse while retaining the existing VLM arch...

Xiuyuan Zhu, Ke Lu, Kun Dong et al. · 0 citations

UvA-DARE (Digital Academic Repository) TWIST & SCOUT: Grounding Multimodal LLM-Experts by Forget-Free Tuning

This paper proposes TWIST, a twin-expert stepwise tuning module that modifies the decoder of the language model using one frozen module pre-trained on image understanding tasks and another learnable one for visual grounding tasks, which allows the MLLM to retain previously learned knowledge and skills, while acquiring...

AritraBhowmik, MohammadMahdiDerakhshani, Dennis C. Koelma et al. · 0 citations
#small language model Preprint Aug 2026

VisLens: Single-Pass Interpretable Visual Search for Multimodal LLMs

VisLens (Visual Focus via Logit Lens), a Visual Search method built on the logit lens, which decodes the semantics held in a hidden state by projecting it through the LLM head, and which matches or exceeds prior baselines while delivering a substantial latency advantage.

Jingfeng He, Sanghwan Kim, Zeynep Akata · 0 citations
Preprint Sep 2026

TempoGround: State-Aware Streaming Visual Grounding with Vision-Language Models

Visual grounding maps language referents to spatial targets and is central to open-vocabulary perception with vision-language models. Existing methods have made substantial progress on single-frame and video-based visual grounding, yet under streaming inputs they still suffer from identity drift, cross-frame inconsiste...

Le-Qian Ding, Jun-Ning Qiu, Man-Wen Yang et al. · 0 citations
#artificial intelligence Preprint Aug 2026

RelCheck: Dual-Evidence Spatial Grounding for VLM Hallucination Correction

Multimodal large language models (MLLMs) fre- quently generate text that is inconsistent with the input image. While object- and attribute-level hallucinations have received considerable attention, relational hallucinations (incorrect de- scriptions of spatial or interactive relationships between objects) remain largel...

Siddhi Patil, N. Saxena, William B. Andreopoulos · 0 citations
Preprint Aug 2026

MLLMCLIP: Feature-Level Distillation of MLLM for Robust Vision-Language Representations

Compared to prior CLIP-enhancement methods, MLLMCLIP achieves state-of-the-art compositional accuracy while delivering consistent gains on standard zero-shot classification and image-text retrieval, showing that feature-level distillation strengthens both compositional and general vision-language representation capabil...

Jongsuk Kim, Qiyu Wu, Zhuoyuan Mao et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.