Skip to content

Category

computer vision

3,022 papers

#artificial intelligence Preprint Sep 2026

FocusVTC: Efficient and High-Performance Visual Text Compression with Adaptive Resolution

Long-context reasoning in large language models incurs substantial computation and memory costs. Visual text compression (VTC) reduces input length by rendering text as images, but fixed-resolution rendering creates a compression-performance trade-off: low DPI saves tokens at the expense of legibility, whereas high DPI...

Fang-Zhi Zhong, Xue-Rui Qiu, Yu-Qi Pan et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Scaling Video Generation for Reasoning: At What Cost?

Symbolic state supervision raises the 20M model's frame accuracy from 31.1% to 67.3% at the same training-data budget, suggesting that learning representations of state changes can complement scaling.

Wei-Hang Guo, Xiao-Yu Wu, Yi-Fei Wang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

ThinkingGuard: Decoding Implicit Hazards via Step-by-Step Risk Attribution in Multimodal Large Language Models

While Multimodal Large Language Models (MLLMs) are increasingly deployed in safety-critical domains, their reliability is threatened by multimodal implicit risks. Unlike explicit threats, these hazards emerge when individually benign text and neutral visual entities logically converge to induce unsafe outputs. Current...

Ruo-Chen Zhang, Yao Huang, Yi-Tong Sun et al. · 0 citations
#artificial intelligence Preprint Sep 2026

FM-ReID: Selective Competitive Token Routing for Object Re-Identification

FM-ReID is proposed, an end-to-end framework that formulates local representation learning as selective competitive token routing as an effective way to augment holistic foundation-model representations, supporting competitive token routing as an effective way to augment holistic foundation-model representations.

Zhi-Qiang Li, Xiao-Wei Zhou, Ze-Yuan Sun et al. · 0 citations
#artificial intelligence Preprint Sep 2026

How Medical VLMs Underutilize Their Vision Encoders: A Dermatology Perspective

A mechanistic analysis of the VLM model's internal attention patterns shows that a simple describe-then-decide prompting strategy increases vision attention by 30-40% during generation, and task-specific fine-tuning improves dermatology classification but reduces cross-domain medical question-answering performance in t...

Janet Wang, Yun-Bei Zhang, Xiao Wang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Similar Choices, Different Attention: Cross-Modal Associations in Humans and Vision-Language Models

Cross-modal associations are systematic pairings of features across modalities, such as the association of'bouba'with round shapes and'kiki'with sharp shapes. Prior work has compared humans and vision-language models (VLMs) on such associations, but often using different stimuli or tasks between humans and models. Here...

Su-Min Hong, Katsumi Ibaraki, Renee Shi et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Staircase Policy: Streaming Inference for World-Action Models with Large Action Chunks

STAIRCASE POLICY is introduced, a streaming inference and training framework that turns a flow-matching VLA into a JEPA-style WAM and partitions a large action chunk into sub-chunks at staggered denoising stages, enabling long-horizon execution without repeated full policy inference.

Guoheng Sun, Chen Chen, Jin Wang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Online Versatile Incremental Learning: Towards Class and Domain-Agnostic Adaptation at Any Time

This work introduces Online VIL (Online Versatile Incremental Learning), a novel scenario where class concepts and visual domains evolve simultaneously online without explicit boundaries, and proposes a novel framework TopFlow, Topology preservation with Flow matching representation that contains two complementary mech...

Jae-Ho Lee, Minji Park, Jun-Yeong Moon et al. · 0 citations
#artificial intelligence Preprint Sep 2026

FineART: Fine-Grained Annotated Robotic Trajectory Dataset and Vision-Language-Action Model for Bimanual Manipulation

FineART is presented, a densely annotated bimanual manipulation dataset comprising 40,543 episodes and 533,913 subtasks across 151 tasks and FineART-VLA is introduced, a vision-language-action policy that predicts its own next subtask to guide its actions.

Jade Choghari, Pepijn Kooijmans, Mansi Agarwal et al. · 0 citations
#artificial intelligence Preprint Sep 2026

AdaKerNet: Neural Kernel Decoding for Task-Adaptive Prediction with Multimodal Large Models

Large foundation models have been introduced with the promise of efficient adaptation to downstream tasks. Yet, under limited supervision, MLLMs, an important class of large foundation models, remain challenging to adapt to various downstream tasks. Adaptation typically relies either on MLLM parameter fine-tuning or on...

Konstantinos D. Polyzos, Eleni Oikonomou, Tara Javidi · 0 citations
#artificial intelligence Preprint Sep 2026

Compress to Remember: Learning Compact Memory via On-Policy Distillation for Long Video Generation

Standard video generators do not natively compact historical context into reusable memory tokens. As generation continues, the growing history makes it increasingly difficult to retain information from earlier frames due to long-context degradation. Key-frame-based approaches address this challenge by retaining selecte...

Xiao-Yu Wu, Wei-Hang Guo, Yi-Fei Wang et al. · 0 citations

From tech blogs

See all →
Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.