Skip to content

Category

computer vision

3,022 papers

#artificial intelligence Preprint Open access Oct 2026

MEND: Label-Free Detection, Localisation, and Correction of Latent Hallucination in World Models

World Models are appearing as the next major frontier in computer vision. However, their robustness is currently largely unexplored. We identify the phenomenon of hallucination in latent World Models: given a state and an action, the predicted next latent can decode to a scene that never occurs. Because the prediction...

Ali J Alrasheed, Aryan Yazdan Parast, Basim Azam et al. · 0 citations
#artificial intelligence Preprint Sep 2026

TED:Text-Axis Evidence Decomposition for Prompted Anomaly Localization

CLIP is a powerful vision-language model, but it was not designed for fine-grained defect localization; CLIP-based anomaly detectors therefore adapt it with prompts or lightweight modules to increase defect sensitivity. We show that stronger sensitivity does not necessarily make local evidence reliable: under domain sh...

Jinyoung Kim, Geonho Kim, Gijeong Park et al. · 0 citations
#artificial intelligence Preprint Sep 2026

On the Relaxation of Conditional Independence Assumption for Image Segmentation

In semantic segmentation, a recent line of RankSEG methods directly optimizes Dice/IoU scores at inference time, improving alignment with evaluation metrics without modifying model training. Despite its theoretical and empirical success, RankSEG relies on the restrictive Conditional Independence Assumption (CIA), which...

Zi-Xun Wang, Ben Dai · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Where MLLMs Fail and Why: Causal Task Decomposition for Capability Failure Diagnosis

End-to-end accuracy on compositional tasks records how often MLLMs fail, but cannot distinguish whether a failure reflects an intrinsic deficit in the targeted capability or a cascading error from an upstream prerequisite. We propose a causal decomposition framework that isolates these two failure modes through control...

Xia Hu, Brian Potetz, Chun-Ta Lu et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Future Video Generation Better Aligns with the Human Visual Cortex than Observed Video

Studying the alignment between the internal representations of vision models and the responses of the visual cortex to the same observed visual stimuli has enabled us to better understand human visual processing. However, studies so far have largely overlooked the fact that the human brain not only processes observed v...

Chang-Bae Bang, Hyungjin Chung, Byung-Hoon Kim · 0 citations
#artificial intelligence Preprint Open access Oct 2026

CRAFT: Causal Responsibility and Failure Tracing in Medical Vision Language Models

As vision language models are increasingly deployed in clinical diagnosis, understanding how they internally resolve competing visual and textual signals becomes a safety imperative. Existing mechanistic analyses remain confined to unimodal text and offer no explanation for why a single misleading sentence can override...

Chunzheng Zhu, Jiaqi Zeng, Hongbo Zhao et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Soft Spatial Reasoning

Large Vision-Language Models (LVLMs) commonly perform spatial reasoning through chain-of-thought (CoT), encoding intermediate reasoning as autoregressive sequences of discrete language tokens. Such hard thinking requires committing to a single token at each step, even when the correct spatial interpretation remains unc...

Rafi Ibn Sultan, Md. Sajid Alam Chowdhury, Saleh Zare Zade et al. · 0 citations
#artificial intelligence Preprint Sep 2026

SpatialCORE: Confidence-Aware Grounded Spatial Reasoning in Large Vision--Language Models

Large Vision-Language Models (LVLMs) have made remarkable progress across visual perception tasks, yet spatial reasoning remains a persistent weakness, especially for questions that require reasoning over visual space. Recent spatial-reasoning methods incorporate generated grounding, where models predict bounding boxes...

Rafi Ibn Sultan, Xiang-Yu Zhou, Mohammad O. S. Chowdhury et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Unveiling the Value of Motion for Cinematic Camera Trajectories

Cinematic camera motion is a fundamental storytelling tool, defined not only by where the camera is positioned in the scene, but also by how it moves in terms of direction and speed. Recent work on camera trajectory generation and alignment to text relies on pose-centric representations. While in principle a network co...

Zi-Qi Zhou, Yu-Jian Yuan, Laura Sevilla-Lara · 0 citations
#artificial intelligence Preprint Open access Oct 2026

ReGain: Restoring Subject Fidelity in Personalization on Synthetic Images

Text-to-image diffusion models are personalized to a subject by DreamBooth fine-tuning on a handful of its images. Increasingly, these images come from a diffusion model rather than a camera. We show that fine-tuning on such synthetic images degrades subject fidelity, producing oversaturated color and excess high-frequ...

Shubhang Bhatnagar, Ishan Bhatnagar, Viraj Shah et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Breaking Babel: A Self-Evolving Multi-Agent System for Long-Form Subtitle Translation

Long-form subtitle translation requires reasoning over discourse and cultural context spanning episodes or entire series, while maintaining consistent terminology and style. Existing single-LLM methods are largely sentence-level, and multi-agent systems often use static workflows that do not adapt to scene complexity o...

Hai-Bo Jin, Xin-Jie Li, N. Sadoughi et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Template-Search Domain Adaptation via Multi-Stage Feature Alignment for Cross-Modal Object Tracking

Visual object tracking typically assumes that the initial template and subsequent search frames share the same sensing modality. In practice, sensor availability or operation may change over time, creating a substantial representation gap between template and search frames. Unlike conventional multi-modal tracking wher...

Fereshteh Aghaee Meibodi, Amir Mehdi Soufi Enayati, Shadi Alijani et al. · 0 citations

From tech blogs

See all →
Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.