Skip to content

Category

computer vision

3,022 papers

#computer vision Preprint Sep 2026

Imagine3D-LLM: Teaching MLLMs to Imagine 3D Scenes Before Answering

Reasoning about the 3D world from multi-view images remains a fundamental challenge for Multimodal Large Language Models (MLLMs). While modern MLLMs handle single-image inputs effectively, they struggle to integrate evidence across viewpoints into a coherent 3D understanding. A growing body of work attempts to close th...

Jaewoo Jung, Hyeonseo Yu, Honggyu An et al. · 0 citations
#computer vision Preprint Sep 2026

VehicleArena: A Realistic Urban Environment for Multi-Agent Driving

Real-world embodied agents often pursue independent objectives within a shared physical environment, where their actions can alter the conditions faced by others. Existing benchmarks, however, typically assume shared goals or explicitly prescribed interaction protocols, leaving such emergent physical coupling underexpl...

Jie Yang, Jia-Jun Chen, Jia-Zheng Zhou et al. · 0 citations
#computer vision Preprint Open access Sep 2026

From Routing Signals to Selective Review: Visual regrounding in MoE VLMs

Vision-language models (VLMs) may accept false visual premises, answering questions about a target object's color, count, location, or state even when it is absent. We call this reliability-critical behavior a target-absence grounding failure. Existing visual-grounding detectors primarily rely on generated responses, h...

Hongzhu Guo, Mohsen Fayyaz, Nanyun Peng · 0 citations
#computer vision Preprint Sep 2026

Chinese-Jev: Bringing System One Model to Chinese-Language Tasks

System One models such as Jev offer an efficient alternative to generative language models for tasks that require decisions rather than open-ended responses. However, existing Jev models exhibit limited Chinese-language decision accuracy, restricting their utility in both general and specialized settings. In this paper...

Ze-Xiao Wang, Zi-Hao Zhang, Xu-Dong Wang et al. · 0 citations
#computer vision Preprint Open access Sep 2026

RAEGNet: Relation-Aware Evidence Graph Network for Harm-Aware Multimodal Fake News Detection

Existing multimodal fake news detection methods often introduce external information to assist detection. However, most of them rely on entity-level retrieval and are therefore prone to introducing event-irrelevant noise. Meanwhile, existing methods mainly focus on improving overall performance and do not account for d...

Wenbin Shen, Guoxuan Qin, Guangxu Yao et al. · 0 citations
#computer vision Preprint Sep 2026

Rethinking Multimodal Fake News Detection in the Generative AI Era

Generative content is increasingly entering the production and dissemination of news, transforming fake news from manually fabricated or simply manipulated material into complex forms in which native and generated content jointly participate. Existing multimodal fake news detection research primarily focuses on veracit...

Wen-Bin Shen, Guo-Xuan Qin, Guang-Xu Yao et al. · 0 citations
#machine learning Preprint Open access Sep 2026

SCAMP: Sparse-anchor Control is One Small Projection

Authoring with a text-to-motion generator needs sparse anchors: chosen joints, at chosen frames, at given positions. Meeting them currently costs a conditioning branch trained for the task, or hundreds of per-clip optimisation steps in the architecture's native variables. In any generator that decodes a continuous...

Pengcheng Fang, Tengjiao Sun, Xiaoyu Zhan et al. · 0 citations
#machine learning Preprint Open access Sep 2026

Vision Wormhole: Latent-Space Communication in Heterogeneous Multi-Agent Systems

Heterogeneous multi-agent systems combine models with different capabilities through a common communication interface. Exchanging internal states directly requires translating between model-specific representations and controlling intermediate computation. We introduce the Vision Wormhole, which repurposes the visual i...

Xiaoze Liu, Ruowang Zhang, Weichen Yu et al. · 0 citations
#machine learning Preprint Open access Sep 2026

Constructive Distortion: Improving MLLMs with Attention-Guided Image Warping

Multimodal large language models (MLLMs) often miss small details and spatial relations in cluttered scenes, leading to errors in fine-grained perceptual grounding. We introduce AttWarp, a lightweight method that allocates more resolution to query-relevant content while compressing less informative areas, all while pre...

Dwip Dalal, Gautam Vashishtha, Utkarsh Mishra et al. · 0 citations
#machine learning Preprint Open access Sep 2026

TeD-Loc: Text Distillation for Weakly Supervised Object Localization

Weakly supervised object localization (WSOL) models can predict both the object class and the spatial regions corresponding to the object, without requiring explicit bounding-box annotations. Given their reliance on classification objectives, traditional WSOL methods, like class activation mapping, tend to focus on the...

Shakeeb Murtaza, Soufiane Belharbi, Alexis Guichemerre et al. · 0 citations
#machine learning Preprint Open access Sep 2026

Beyond Missing Rates: Rethinking Incomplete Multi-View Clustering with Protocol Divergence

Incomplete multi-view clustering (IMVC) is typically evaluated by retraining separate models under different missing-view configurations. Evaluations indexed only by nominal missing rate can overlook differences in observation structure across missing-view protocols. We show that missing-data protocols with identical n...

Haolu Liu, Xiyue Wang, Xuanting Xie et al. · 0 citations
#machine learning Preprint Open access Sep 2026

Formalizing the Sampling Design Space of Diffusion-Based Generative Models via Adaptive Solvers and Wasserstein-Bounded Timesteps

Diffusion-based generative models have achieved remarkable performance across various domains, yet their practical deployment is often limited by high sampling costs. While prior work focuses on training objectives or individual solvers, the broader sampling design problem, specifically solver selection and scheduling,...

Sangwoo Jo, Sungjoon Choi · 0 citations

From tech blogs

See all →
Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.