Skip to content

Category

computer vision

2,913 papers

#artificial intelligence Preprint Open access Oct 2026

From Pixel to Coding: Evaluating the Figure Reproduction Capabilities of MLLMs

Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities in both visual understanding and code generation. However, existing benchmarks typically evaluate these two modalities in isolation, lacking a dedicated assessment of their unification, i.e., how a model can perceive complex visual struc...

Zijian Chen, Zhengyu Chen, Bohan Liang et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

UltraText Bench: A Comprehensive Bilingual Benchmark for Evaluating Visual Text Rendering in Image Generation

Dense visual text requires image generators to reproduce long strings across multiple regions with correct placement and legibility. As short-string rendering improves, evaluation must test sustained performance across more demanding scenes. We introduce UltraText Bench, a bilingual benchmark for prompt-only generation...

Deyuan Liu, Yihao Hu, Jingxuan Zhang et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

DisParQ: Self-Supervised Part Concepts for Interpretable Vision Foundation Models

Concept-based vision models represent images through an intermediate layer of human-inspectable concepts, so what a model relies on can be traced to those concepts. However, those models are often limited to fixed categories or depend on language to define their concepts. We introduce DisParQ (Discrete Parts with Quant...

Adam Pardyl, Siddhartha Gairola, Sukrut Rao et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

PARC-Loc: Text-to-Point-Cloud Localization with Partial Assignment and Relational Consistency

Text-to-point-cloud localization estimates a position in a city-scale 3D map from descriptions of surrounding objects. Existing coarse-to-fine methods retrieve submaps using aggregate learned compatibility and then localize within a selected submap. However, repetitive or similar urban objects can inflate the embedding...

Shengkai Ma, Zhenyu Hou, Weihua Cao · 0 citations
#artificial intelligence Preprint Open access Oct 2026

What Makes Synthetic Hard Negatives Work in Vision-Language Pretraining?

Synthetic hard negatives generated in the representation space have proven effective for unimodal self-supervised learning, but transferring this idea to vision-language pretraining is not straightforward. We analyze six representation-space synthesis strategies and identify two failure modes in their transfer to visio...

Nikos Giakoumoglou, Paschalis Giakoumoglou, Andreas Floros et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

WAPR: A Foundation Model for Wide-Angle Refinement in Unseen Object Pose Estimation

Real-world applications require 6D pose estimation to be accurate, fast, and scalable to unseen objects. This paper introduces WAPR, a zero-shot wide-angle pose refinement model that refines candidate poses with rotational deviations up to 90 degrees. With as few as 12 candidate poses per detected object instance, WAPR...

Yulin Wang, Mengting Hu, Hongli Li et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

KASALv2: Fully Automatic 3D Rotational Symmetry Classification and Axis Localization

Rotational symmetry is an important prior in 6D pose estimation, improving pose accuracy and supporting symmetry-aware evaluation. However, current symmetry annotations for 3D objects remain largely manual or semi-automatic, often requiring predefined types or orders, which limits scalability. This work introduces a fu...

Mengxin Zhang, Yulin Wang, Chen Luo et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

OmniCam: Omni-Camera Trajectory Generation via Geometry-Grounded Pose Token Learning

Camera trajectories control viewpoint changes in video generation, scene reconstruction, and robotic perception. Generating them from language requires both scene geometry and target-aware framing. We introduce OmniCam, an autoregressive model that generates camera pose sequences from a single panorama and textual traj...

Zhenyang Liu, Chenjie Cao, Yisu Zhang et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

SpatialUQ: Post-Hoc Uncertainty Quantification from Spatial Consistency in Black-Box Vision Models

Clinical vision models are often deployed as frozen black boxes with no access to internals, retraining, or ground truth at inference time. We introduce \textbf{SpatialUQ}, a post-hoc uncertainty method using only output probabilities. It measures the Jensen-Shannon divergence between the global prediction and the mean...

Md Kawsher Mahbub, Milon Biswas, Mirza Niaz Morshed et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Iris-3B: Going Beyond the Latent with Pixel-Space Diffusion Training, Conversion and Fine-Tuning

Pixel-space diffusion models avoid the lossy VAE of latent models, which suggests an advantage on downstream tasks where fine-grained detail matters. We test this claim along both routes to a pixel-space backbone. We pretrain Iris-3B, a 3B-parameter pixel-space text-to-image transformer, from scratch through a $256\to5...

Hanqiu Li Cai (SperidLabs), Chema Garabito (SperidLabs) · 0 citations
#artificial intelligence Preprint Oct 2026

Mixture of Layers: Dynamic Layer Routing for Visual Reasoning

Pre-trained vision encoders contain layer-wise visual representations that differ in spatial granularity, semantic abstraction, and sensitivity to local details. However, most Multimodal Large Language Models (MLLMs) rely on only the final or penultimate vision encoder representations or fixed aggregation rules, making...

Jeonghwan Kim, S. Stoica, Ji-Wan Chung et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Shared Geometry As A Rosetta Stone: Cross-Modal Alignment Without Paired Data

Multimodal representations enable zero-shot classification and retrieval, but aligning independently trained models usually requires large amounts of paired data. Yet, the Platonic Representation Hypothesis suggests that models trained on different modalities may converge spontaneously toward a shared representation ge...

Dominik Schnaus, Thomas Dag\`es, Daniel Cremers et al. · 0 citations

From tech blogs

See all →
Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.