Skip to content

Author

Salman H. Khan

3 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Oct 2026

HeiCo-FOCUS: A Clinically Grounded Dataset for Long-Context Video Understanding

Recent advances in Vision-Language Models (VLMs) have led to rapid progress in video understanding across a wide range of benchmark tasks. However, existing evaluations largely focus on short-term reasoning, failing to assess a critical capability: maintaining cumulative temporal consistency over extended time horizons...

Leon D. Mayer, Lucas Luttner, P. Godau et al. · 0 citations
Preprint Aug 2026

Locate Anything in Videos: Rethinking Efficient Generative Spatio-Temporal Video Grounding

Spatio-temporal video grounding (STVG) requires models to identify when a referred event occurs and localize the target entity throughout that interval. Existing multimodal large language models typically serialize dense localization trajectories autoregressively, causing decoding latency to grow with tube length and a...

H. Rasheed, H. Siddiqui, Ming-Hsuan Yang et al. · 0 citations
Preprint Aug 2026

Training-Free Speech-Centric Omni Understanding with Frozen VLMs

Audio-visual understanding remains challenging because models must jointly interpret spoken content, visual events, and their temporal relationships. Existing omni models typically introduce dedicated audio encoders and rely on expensive audio-video-text training, tightly coupling omni capability to specific VLM backbo...

Ankan Deria, H. Rasheed, Xi-Lin He et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.