Skip to content

Category

computer vision

3,022 papers

#machine learning Preprint Open access Oct 2026

TEMPEST: Temporal Embeddings for Scalable Driver Identification via Angular Margin Learning

Scalable driver identification requires embedding models that maintain discriminative performance as fleet size grows, yet existing triplet-loss formulations degrade rapidly with driver pool size and overfit to session-specific patterns under rigorous temporal evaluation. We introduce TEMPEST, a Temporal Convolutional...

Kyle Musgrove, Dylan B. Lewis, Sarah Powers et al. · 0 citations
#computer vision Preprint Open access Oct 2026

Answer with Evidence: Consistency-Aware Grounded Visual Question Answering for Roadside Traffic Scenes

Roadside traffic reasoning requires every free-form textual claim to be backed by visual evidence. Existing grounded multimodal large language models (MLLMs) frequently exhibit say-point mismatch, in which the textual answer contradicts the bounding boxes the model localizes. Evaluation metrics that score answers and b...

Runwei Guan, Rongsheng Hu, Shangshu Chen et al. · 0 citations
#computer vision Preprint Open access Oct 2026

Beyond the Good, the Bad, and the Ugly: Colormap Assessment through Data-Aware Perceptual Metric

Continuous colormaps are widely used to visualize scalar fields, and their quality is typically evaluated using measures such as discriminative power and uniformity. Existing measures primarily characterize the intrinsic perceptual properties of the colormap itself, largely independent of the underlying data distributi...

Xi Duan, Yiwei Lin, Shiqing Xin et al. · 0 citations
#computer vision Preprint Open access Oct 2026

EvoDesign: Agentic Editable Diagram Creation via Design Expertise Evolution

High-fidelity diagram creation requires the complex orchestration of semantic topology, visual styling, and spatial layout, posing a significant challenge for automated systems. Existing methods also suffer from a representation gap: pixel-based models often lack precise control, while code-based synthesis limits intui...

Tianfu Wang, Leilei Ding, Ziyang Tao et al. · 0 citations
#computer vision Preprint Open access Oct 2026

LinguDistill: Recovering Linguistic Ability in Vision-Language Models via Selective Cross-Modal Distillation

Turning a pretrained language model (LM) into a vision-language model (VLM) through multimodal fine-tuning often erodes its native language ability, a form of catastrophic forgetting that shows up even on text-only tasks. This loss is hard to undo with further fine-tuning, and existing remedies add adapters or alignmen...

Patrick Amadeus Irawan, Erland Hilman Fuadi, Shanu Kumar et al. · 0 citations
#computer vision Preprint Open access Oct 2026

3ViewSense: Spatial and Mental Perspective Reasoning from Orthographic Views in Vision-Language Models

Current Large Language Models have achieved Olympiad-level logic, yet Vision-Language Models paradoxically falter on elementary spatial tasks like block counting. This capability mismatch reveals a critical ``spatial intelligence gap,'' where models fail to construct coherent 3D mental representations from 2D observati...

Shaoxiong Zhan, Yanlin Lai, Zheng Liu et al. · 0 citations
#computer vision Preprint Open access Oct 2026

MMLongCite: A Benchmark for Evaluating Faithfulness of Long-Context Vision-Language Models

The rapid advancement of long-context vision language models (LCVLMs) has led to a significant expansion of their context windows. However, an extended context window does not guarantee the effective utilization of the context, posing a critical challenge for real-world applications. Current evaluations of such long-co...

Keyan Zhou, Zecheng Tang, Lingfeng Ming et al. · 0 citations
#computer vision Preprint Open access Oct 2026

DEFAME: Dynamic Evidence-based FAct-checking with Multimodal Experts

The proliferation of disinformation demands reliable and scalable fact-checking solutions. We present Dynamic Evidence-based FAct-checking with Multimodal Experts (DEFAME), a modular, zero-shot MLLM pipeline for open-domain, text-image claim verification. DEFAME operates in a six-stage process, dynamically selecting th...

Tobias Braun, Mark Rothermel, Marcus Rohrbach et al. · 0 citations
#computer vision Preprint Open access Oct 2026

AVOC: Enhancing Hour-Level Audio-Video Understanding in Omni-Modal LLMs via Retrieval-Inspired Token Compression

Multimodal Large Language Models have achieved remarkable progress in short-form audio-video understanding, yet long-form audio-video comprehension remains challenged by limited context windows and severe information redundancy. To address these bottlenecks, we propose AVOC, a framework for long-form audio-video unders...

Yijing Chen, Wenhui Tan, Xiaoyi Yu et al. · 0 citations
#computer vision Preprint Oct 2026

Efficient Test-time Adaptation through Candidate Verification and Divergence Shifts

Vision-language models (VLMs) achieve strong zero-shot transferability but remain vulnerable to target-domain shifts at inference time. Test-time adaptation (TTA) offers a practical remedy, yet most existing VLM-TTA methods follow a prediction-side adaptation paradigm. They use test samples to adjust logits, prototypes...

Seungmin Oh, Seung-Hun Kang, Jongbin Ryu · 0 citations
#computer vision Preprint Open access Oct 2026

PlotGround: Grounding Plot Digitization in Real Scientific Figures and Their Source Data

Scientific figures often encode quantitative results that are not readily available in machine-readable form, making accurate plot digitization important for verifying and reusing published findings. Yet it remains unclear how accurately current models recover plotted values from real scientific figures, as existing be...

Yaohui Zhang, Binxu Li, Haoyi Duan et al. · 0 citations
#computer vision Preprint Open access Oct 2026

Atomic Visual Entailment: Enhancing Zero-Shot Vision-Language Reasoning through Atomic Fact Decomposition and Learned Selection

Visual entailment (VE) asks whether an image supports, contradicts, or leaves undecided a textual hypothesis. Strong results come from fine-tuning large vision-language models on labelled data, while zero-shot and hybrid approaches remain far behind. A VE hypothesis often bundles several visual claims, yet existing zer...

Nallathambi Vethiappan, Derya Soydaner, Gijs Wijnholds · 0 citations

From tech blogs

See all →
Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.