Skip to content

Category

computer vision

3,022 papers

#artificial intelligence Preprint Open access Oct 2026

ReMAP: Restoring the Perceptual Cycle with Reasoning-Time Latent Visual Memory

As multimodal large language models (MLLMs) reason for longer, attention to the initial visual input diminishes, weakening visual grounding. Visual memory reintroduces visual evidence during reasoning. We conduct a controlled analysis of visual memory along three axes: curation, organization, and access. We find that l...

Hao Jiang, Zhanyu Guo, Chenwei Wu et al. · 0 citations
#artificial intelligence Preprint Oct 2026

Enhancing Long-Video VLM Embeddings with Query-Aware Streaming Latent Reasoning

Long-video embedding requires capturing sparse query-relevant evidence under a limited visual-token budget. Uniform sampling can miss brief events in videos spanning minutes or hours, whereas encoding more frames in a single context increases memory and computation. We introduce \textbf{Query-Aware Streaming Latent Rea...

Hao-Zhe Chi, Song Jin, Yang Jin et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Active-DiNTS: Active Differentiable Network Topology Search

Neural Architecture Search (NAS) has proved to be a strong alternative to manual network design, but applying it to 3D medical image segmentation is limited by two well-known costs, large annotation budgets and multi-GPU clusters. Thus, this paper introduces Active-DiNTS (Active Differentiable Network Topology Search),...

Gean Trindade Pereira, Thierry Urruty, Muriel Visani et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

ARISE: Adaptive Agentic Reasoning with Image-grounded Self-Evaluation for Interpretable IBD Assessment

Inflammatory bowel disease (IBD) requires frequent imaging-based assessment, yet interpretation of modalities such as wireless capsule endoscopy (WCE) and intestinal ultrasound remains heavily dependent on specialist expertise. Vision-Language Models (VLMs) demonstrate significant potential in multimodal medical image...

Pronoma Banerjee, Anuva Shah, Jason Wu et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Knossos and Ariadne: Benchmarking and Learning Complete Diagram Topology Extraction with Vision-Language Models

Structural diagrams are widely used to represent complex systems and relational information across scientific, engineering, procedural, and spatial domains. Recent vision-language models (VLMs) have become increasingly capable of recognizing diagram elements and reasoning about their content, while complete diagram top...

Bangwei Guo, Xujiang Zhao, Shengyu Chen et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

ConEx: Human-Interpretable Saliency Maps via Concept-Aware Attribution

Many visual explanation methods in computer vision highlight pixel importance but struggle to link these low-level cues to semantically meaningful concepts, limiting their interpretability and trustworthiness. We introduce Concept-based Explanations (ConEx), a novel framework that bridges saliency visualization with co...

Yehonatan Elisha, Oren Barkan, Ziv Weiss Haddad et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Cross-Modal Solar Image Synthesis: Adapting the Surya Foundation Model from He I 10830 {\AA} to EUV Translation and Coronal Hole Segmentation

The long observational record of He I 10830 {\AA} offers a means to investigate solar morphology before modern extreme-ultraviolet (EUV) imaging. We adapt the Surya solar foundation model to predict Solar Dynamics Observatory/Atmospheric Imaging Assembly (SDO/AIA) 94, 193, and 304 {\AA} images and a coronal hole (CH) p...

Marco Marena, Andr\'es Mu\~noz Jaramillo, Qin Li et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

EgoExo-Next:Benchmarking Vision-Language Models on Visual-Option Next-State and Cross-View Reasoning

Vision-language models (VLMs) are increasingly evaluated for egocentric and cross-view video reasoning, yet existing benchmarks largely focus on semantic event understanding, temporal relations, or correspondence between already observed views, leaving their ability to reason directly about future visual states underex...

Yutong Li, Molin Wang, Xiaotong Li et al. · 0 citations

Localization Lens for Improving Medical Vision-Language Models

Medical Vision-Language Models (Med-VLMs) have demonstrated strong capabilities in clinical tasks. However, they often struggle to understand anatomical structures and spatial positioning, which are crucial for medical reasoning. To address this, we propose a localization-aware enhancement to the Med-VLM pipeline, intr...

H. Farooq, Murtaza Taj, Mehwish Nasim et al. · 2 citations
#artificial intelligence Preprint Open access Oct 2026

Understanding Clustering in Slot Attention via Particle Dynamics

Studying attention through the lens of interacting particle dynamics has shown how token clustering can emerge from the underlying dynamics. We extend this perspective to slot attention, a method for object-centric image segmentation and representation learning in which learned components obscure how much of the cluste...

Vasudev Joy, Rajat Rasal, Avinash Kori et al. · 0 citations
#artificial intelligence Preprint Oct 2026

Video2World: Benchmarking Coding Agents for Interactive World Modeling from Embodied Videos

Building interactive simulators from real-world observations is a promising way to scale embodied data, but current pipelines still rely heavily on manual environment construction and calibration. We study whether frontier foundation models and coding agents can automate this process end to end. We formulate \emph{auto...

Jin-Zhou Tang, Z. Zhang, Jing Yang et al. · 0 citations
#artificial intelligence Preprint Oct 2026

AgroGround: Multi-Granularity Grounded Recognition in Agriculture

Agricultural visual models are typically evaluated for either recognition or localization, but reliable diagnosis requires identifying what is present and localizing the evidence. Agricultural visual question answering (VQA) datasets carry rich semantic labels but rarely link them to image regions, and adding such anno...

Abdulla Alshehhi, Zong-Yan Han, R. Anwer · 0 citations

From tech blogs

See all →
Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.