Skip to content

Category

computer vision

3,022 papers

#artificial intelligence Preprint Sep 2026

PolyOCR-Venus: Unified OCR Foundation Models for Text-Centric Visual Intelligence

This report presents PolyOCR, a family of unified OCR foundation models of varying scales that combines a shared instruction-following framework with a large-scale data engine that converts heterogeneous visual resources into quality-verified OCR supervision.

Guang-Zhan Huang, Yong-Shuo Zhang, Bing-Tao Fu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

VIF-Bench: Evaluating Visual Instruction Following in Multi-Reference Image Generation

Recent multimodal image generation models can take multiple images and textual instructions as input, enabling reference-based generation guided not only by text but also by visual instructions such as layouts, arrows, and pose cues. However, existing benchmarks do not evaluate the joint setting in which multiple refer...

Yuta Oshima, Masakazu Yoshimura, Masahiro Suzuki et al. · 0 citations
#artificial intelligence Preprint Open access Sep 2026

FedSocket: Recipient-Executable Knowledge Exchange for Heterogeneous Multimodal Federated Learning

Federated knowledge must remain usable by recipients with different modalities, private architectures, and tasks. We present FedSocket, which makes recipient execution a design requirement of the exchanged model. A shared Q combines recipient-computable inputs, task-owned outputs, and ownership-aware aggregation, conne...

Xinyuan Zhao · 0 citations
#artificial intelligence Preprint Sep 2026

TReVS: Integrating Textual Relevance and Visual Saliency for Efficient Vision-Language Model Token Pruning

TReVS is proposed, a training-free framework that combines textual relevance with vision-encoder saliency for pre-LLM pruning and leverages high-variance attention heads to remove task-irrelevant tokens at shallow-to-intermediate layers of the LLM.

Jing Wang, Zhi-Ping Wu, Dong-Dong Ren et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Hierarchical Compression of Vision-Language Model Benchmarks

The analyses show how VLM evaluation behaves as model panels grow and evolve, providing guidance for designing future benchmarks that are more efficient, robust to model turnover, and explicit about the limits of evaluation-side pruning.

Hyunjong Ok, Seung-Gu Kang, Jaeho Lee · 0 citations
#artificial intelligence Preprint Sep 2026

LazySloth: Bounded LLM-based Lazy Tree Search for Fast Long Video Comprehension

Modern vision-language models (VLMs) have shown promising results in long-video understanding due to the rich semantic information they can capture. However, most methods focus on coarse captioning of extracted image frames that are computationally inefficient and require models with large context windows. While past w...

Arka Mukherjee, Kaleen Shrestha, Larissa Zhu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Encore: Few-Shot Agentic Discovery of Manipulation Strategies

Coding agents can now write, run, and debug programs with little human help. Robot tasks, however, are usually specified by a sentence that leaves out how to grasp, in what order to make contact, and what the result should look like, and an agent given only the sentence must find these details by trial and error. We in...

Yi-Fan Kang, Zihan Wang, Zhi-Wen Fan et al. · 0 citations
#artificial intelligence Preprint Open access Sep 2026

TAEC: Trajectory-Aware Evidence Coordination for Multi-Step Visual RAG

Multi-step visual retrieval-augmented generation (RAG) answers complex questions by repeatedly retrieving visual evidence, updating an intermediate state, and deciding whether to continue searching or answer. Yet retrieving relevant evidence does not ensure its effective use throughout the reasoning trajectory. As mult...

Yalun Wu, Bingzhou Wang, Boyang Wang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

VISTA-Bench: Benchmarking Multilingual Image Translation with Image-Specific Rubrics

Image translation is a fundamental capability of multimodal models for multilingual applications, requiring visual understanding and meaning preservation across languages. However, existing benchmarks have limited language coverage and often lack explicit image-specific evaluation criteria, making it difficult to compr...

Bo Lv, Mao Zheng, Zheng Li et al. · 0 citations
#artificial intelligence Preprint Sep 2026

UniAfford: Token-Routed Multitask Learning for Generalizable 2D-3D Affordance Perception

Affordance perception aims to localize actionable regions supporting embodied interaction, yet 2D and 3D affordance grounding have evolved as separate problems, with different task definitions, supervision formats, datasets, and evaluation protocols. This fragmentation limits the learning of transferable object-afforda...

Yu-Hao Liu, Yi-Ming Zhong, Han-Qing Wang et al. · 0 citations

From tech blogs

See all →
Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.