Skip to content

Category

computer vision

3,022 papers

#artificial intelligence Preprint Open access Oct 2026

Cross-Cultural Value Attribution in Large Vision-Language Models

The rapid adoption of large vision-language models (LVLMs) in recent years has been accompanied by growing fairness concerns due to their propensity to reinforce harmful societal stereotypes. While significant attention has been paid to such fairness concerns in the context of social biases, relatively little prior wor...

Phillip Howard, Xin Su, Kathleen C. Fraser · 0 citations
#computer vision Preprint Open access Oct 2026

AVMeme Exam: A Multimodal Multilingual Multicultural Benchmark for LLMs' Contextual and Cultural Knowledge and Thinking

Internet audio-visual clips convey meaning through time-varying sound and motion, which extend beyond what text alone can represent. To examine whether AI models can understand such signals in human cultural contexts, we introduce AVMeme Exam, a human-curated benchmark of over one thousand iconic Internet sounds and vi...

Xilin Jiang, Qiaolin Wang, Junkai Wu et al. · 0 citations
#computer vision Preprint Oct 2026

The Failure Is in the Readout: Fine-Grained Emotion Recognition Benchmarks Measure Elicitation, Not Perception

Fine-grained emotion recognition supports therapy tools and social robots, but it needs facial data, which raises privacy and data-protection concerns. EmoNet-Face-HQ answers that with generated portraits, expert-rated over a $40$-category taxonomy far finer than the usual six to eight basic emotions. Under the protoco...

T. Hallmen, Fabian Deuser, Robin-Nico Kampa et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

VisionWeave: Weaving Elastic Visual Representations as a Native Capability of MLLMs

Multimodal large language models have become the dominant paradigm for visual understanding, but incur substantial costs by encoding inputs into dense, fixed-size patch tokens. However, visual information is unevenly distributed: some regions require fine-grained detail, while others admit compact representations. Down...

Yuan Feng, Qize Yang, Ruizhe Chen et al. · 0 citations
#computer vision Preprint Open access Oct 2026

Tracking Is Not Permanence: What Video World Models Keep of a Hidden Object

Video world models track objects they can see; we ask what they keep of objects they cannot. We hide an object from a frozen V-JEPA 2 predictor and compare its prediction for the hidden region with the encoder's representation of two worlds that differ only inside that region. The predictor's decision keeps a stationar...

Peng Xie, Amr Alanwar · 0 citations
#artificial intelligence Preprint Open access Oct 2026

A doctrine-grounded visual question answering dataset for Tactical Combat Casualty Care

Tactical Combat Casualty Care (TC3) requires responders to connect visual observations of injuries and interventions with established clinical guidance. Developing vision-language models to support this process requires supervision that links visible evidence to traceable doctrine. We present TC3-VQA, a dataset constru...

Junseob Kim, Jade Chng, Ayman Ali et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Smart Content Ingestion for Generative AI Workloads

The evolution of machine learning has progressively changed where intelligence resides in an AI system. In conventional machine learning the task, data representation, labels and model architecture were tightly coupled, so data preparation was narrow, schema-bound and visible. Generative AI decouples the model from any...

Abbas Raza Ali, Muhammad Ajmal Siddiqui, Moona Zahid · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Visual Abstention in Unified Multimodal Models

Unified multimodal models (UMMs) integrate understanding and generation, yet their generative behavior is rarely governed by what they understand about the task. We formalize visual abstention: when a requested visual transformation is impossible under the task's rules, the model should recognize that no valid solution...

Chufan Shi, Cheng Yang, Tiannuo Yang et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

SAE++: Cascaded Sparse Autoencoders Learn Multi-Level Visual Concepts in Multimodal LLMs

Multimodal Large Language Models (MLLMs) have demonstrated strong performance on vision-language tasks, yet their internal visual representations remain difficult to interpret. Sparse Autoencoders (SAEs) provide a scalable way to decompose dense model activations into sparse, interpretable features. However, existing S...

Yusong Zhao, Hengyi Wang, Tanuja Ganu et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Learning Visual Feature-Based World Models via Residual Latent Action

World models predict future transitions from observations and actions. Existing works predominantly focus on image generation only. Visual feature-based world models, on the other hand, predict future visual features instead of raw video pixels, offering a promising alternative that is more efficient and less prone to...

Xinyu Zhang, Zhengtong Xu, Yutian Tao et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Von Neumann Networks

In the mid-twentieth century, mathematician and polymath John von Neumann created a computational system on an array of cells as a simple model of the human brain, where each cell had one of a finite set of roles or states that he predicted would be modelled by a diffusion process. In this work, we show that such a sys...

Shekhar S. Chandra · 0 citations
#machine learning Preprint Open access Oct 2026

Generalizable Dense Reward for Long-Horizon Robotic Tasks

Existing robotic foundation policies are trained primarily via large-scale imitation learning. While such models demonstrate strong capabilities, they often struggle with long-horizon tasks due to distribution shift and error accumulation. While reinforcement learning (RL) can finetune these models, it cannot work well...

Silong Yong, Stephen Sheng, Carl Qi et al. · 0 citations

From tech blogs

See all →
Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.