Skip to content

Category

computer vision

3,022 papers

#artificial intelligence Preprint Open access Oct 2026

UrbanVLA: A Vision-Language-Action Model for Urban Micromobility

Urban micromobility applications, such as delivery robots, demand reliable navigation across large-scale urban environments while following long-horizon route instructions. This task is particularly challenging due to the dynamic and unstructured nature of real-world city areas, yet most existing navigation methods rem...

Anqi Li, Zhiyong Wang, Jiazhao Zhang et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

What Drives Compositional Generalization in Visual Generative Models? The Importance of Continuous Training Objectives

Compositional generalization, the ability to generate novel combinations of known concepts, is a key ingredient for visual generative models. Yet, not all mechanisms that enable or inhibit it are fully understood. In this work, we conduct a systematic study of which design choices critically determine compositional gen...

Karim Farid, Rajat Sahay, Yumna Ali Alnaggar et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

COLORA: Efficient Fine-Tuning for Convolutional Models with a Study Case on Optical Coherence Tomography Image Classification

We introduce CoLoRA (Convolutional Low-Rank Adaptation), a parameter-efficient fine-tuning method for convolutional neural networks (CNNs). CoLoRA extends LoRA to convolutional layers by decomposing kernel updates into lightweight depthwise and pointwise components. This design reduces the number of trainable convoluti...

Mariano Rivera, Angello Hoyos · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Architect-Ant: Editable Automatic Furnishing of Architectural Floor Plans

Furnished floor plans support real-estate visualization, interior design, and architectural workflows, yet automatic furnishing remains challenged by limited real-world data and the need to satisfy interacting geometric and functional constraints. We ask whether professional furnishing knowledge can be learned from rea...

Fedor Rodionov, Aleksandar Cvejic, Michael Birsak et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Spatial Atlas: Compute-Grounded Reasoning for Spatial-Aware Research Agent Benchmarks

We describe compute-grounded reasoning (CGR), a design pattern in which code computes selected sub-problems from explicit intermediate representations before a language model answers. Spatial Atlas implements CGR as an Agent2Agent (A2A) server with a spatial question-answering handler and a machine-learning engineering...

Arun Sharma · 0 citations
#artificial intelligence Preprint Open access Oct 2026

SONIC-O1: A Real-World Benchmark for Evaluating Multimodal Large Language Models on Audio-Video Understanding

Multimodal Large Language Models (MLLMs) are a major focus of recent AI research. However, most prior work focuses on static image understanding, while their ability to process sequential audio-video data remains underexplored. This gap highlights the need for a high-quality benchmark to systematically evaluate MLLM pe...

Ahmed Y. Radwan, Christos Emmanouilidis, Hina Tabassum et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

MMMG: a Comprehensive and Reliable Benchmark for Multitask Multimodal Generation

Automatically evaluating multimodal generation presents a significant challenge, as automated metrics often struggle to align with human evaluation, especially for complex tasks that involve multiple modalities. We present MMMG, the first benchmark to bring the verifiable-task paradigm to multimodal generation, spannin...

Jihan Yao, Yushi Hu, Wenyuan Wang et al. · 0 citations
#artificial intelligence Preprint Oct 2026

One Basis to Animate Them All: Gaussian Blendshape Distillation for Real-Time Avatars

3D Gaussian avatars support fast rendering, however, their real-time animation is often challenged by the costly neural inference. We address this bottleneck and show that the animation of pretrained avatar models can be closely approximated by a linear combination of identity-independent blendshapes. Building on this...

Ramazan Fazylov, Stamatis Lefkimmiatis, Ivan Laptev · 0 citations
#artificial intelligence Preprint Oct 2026

SILSA: Sliding-Window Slice Latents for Topology-Preserving High-Resolution 3D Generation

High-resolution 3D generation increasingly relies on voxel latents and multi-stage pipelines that first predict active structure and then synthesize local geometry. While effective, this design fragments continuous surfaces into many local tokens, inflates generation cost, and often weakens topological consistency for...

Tianjiao Yu, Xin-Zhuo Li, Yi-Fan Shen et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

DMAD: Distribution Matching as Adversarial Distillation for Fast Visual Generation

Distribution Matching Distillation (DMD) trains a few-step student from the difference between separately estimated target and student scores, so it must keep an auxiliary diffusion model fitted to the student's evolving distribution at extra memory and computation cost. We introduce DMAD, Distribution Matching as Adve...

Zhengming Yu, Junkun Yuan, Haotian Yang et al. · 0 citations
#artificial intelligence Preprint Oct 2026

Generative Cinematographer: Composing Camera and Object Motion in 3D

Current controllable video generation systems often rely on 2D motion trajectories or sparse drag signals for object motion. These controls are ambiguous because the same 2D trajectory can correspond to different 3D motions, especially when the camera and objects move simultaneously. We present Generative Cinematograph...

Jia-Han Zhang, Chao-Hao Yang, N. Guruprasad et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

MIRTO: a registration-gated, multiverse-tested evaluation protocol for unsupervised anomaly segmentation in brain MRI

Unsupervised anomaly detection (UAD) methods for brain MRI are ranked by a single score, yet that score rests on choices that are rarely reported: how each anomaly map is aligned with the reference, how and on which data the threshold is set, and which false-positive budget, metric, aggregation and lesion definition ar...

Negin Kafee Hernashki, Soumick Chatterjee · 0 citations

From tech blogs

See all →
Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.