Skip to content

Author

F. M. Thoker

We have 2 of 17 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Sep 2026

Composition, Not Conversation: VLMs Lose the Scene, Not the Thread

Vision-language models (VLMs) increasingly reason over visual evidence that is cropped, segmented, retrieved, or revealed over time. Yet most VQA benchmarks present the complete image and question at once. We ask what models lose when the same information is fragmented. We introduce Layered-VQA, with 93 scenes and 300...

L. D. M. S. Sai Teja, Ufaq Khan, N. S. G. Krishna et al. · 0 citations
#small language model Preprint Sep 2026

Advancing Video-Text Pretraining with Multi-View Captions

This work proposes a large-scale multimodal large language model-based supervision generation framework that improves supervision diversity, fidelity, and semantic coverage, and introduces a granularity-aware text representation with separate CLS tokens for summary and detailed views.

F. M. Thoker, Renaud Vandeghen, Karen Sanchez et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.