Skip to content

Category

computer vision

3,022 papers

#artificial intelligence Preprint Sep 2026

Increasing Width Allows Greedy Layer-wise Training to Rival End-to-End Backpropagation in Self-Supervised Learning

End-to-end backpropagation has been the dominant mode of training in deep learning, allowing for the coordination of parameter updates across layers of a neural network. Prior studies have explored alternative -- and, in some cases, simpler -- training mechanisms, showing that they can sometimes achieve performance sim...

Syon Mansur, Joel Zylberberg · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Personalized Image Generation with Reasoning and Reflection

Personalized image generation has remained narrowly focused on conditional synthesis from curated visual exemplars, rather than capturing who a user is. In practice, however, a user's personal context is much richer, comprising reviews, posts, images, captions, and metadata accumulated over time. A truly personalized g...

Bo Ni, Ngoc N. Tran, Qinwen Ge et al. · 0 citations
#artificial intelligence Review Sep 2026

VisionQ: VLM-as-a-Judge Taxonomy, Dataset and Benchmark for Qualitative Analysis in Computer Vision

Qualitative comparison figures are central evidence in computer vision papers, and vision-language models (VLMs) are increasingly used to judge them. Yet existing benchmarks score only scalar quality or overall preference, so a judge can be rewarded for picking the preferred image for the wrong visual reason. We introd...

Xuân Định Vũ, Duc-Hai Nguyen, M. Dao et al. · 0 citations
#artificial intelligence Preprint Sep 2026

PACT: End-to-End Learning of Human Pose, Contacts, and Forces from Video

Human motion, environmental contacts, and interaction forces are governed by common physical laws, yet existing approaches typically separate visual pose reconstruction from contact and force estimation. This separation limits joint reasoning and can propagate errors between stages. We introduce PACT, an end-to-end mod...

Rikhat Akizhanov, Yang-Song Zhang, Nikolai Kaliazin et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Scores That Hold, Benchmarks That Leak: Measuring Dataset Contamination in Public Brain-Tumor MRI Classification

Automated classification of brain tumors from MRI is a heavily published application of deep learning in medical imaging, with reported accuracies on public benchmarks routinely exceeding 98%. However, accuracy does not capture a critical dimension of benchmark quality: dataset integrity, defined as the independence of...

Bhanu Prakash Vangala, Sowmya Guda, Latha Peddi et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Manifold-Constrained Initial Noise Optimization for Efficient Generative Model Alignment

Recent advances in distillation and flow-map models have enabled deterministic one- or few-step generation for high-quality data, facilitating a new branch of reward alignment approaches that directly optimize the initial noise from a Gaussian distribution. However, most existing initial-noise optimization methods rely...

Jin-Ho Chang, Jong Chul Ye · 0 citations
#artificial intelligence Preprint Sep 2026

Diffusion Editing with Soft Mask: Pixel Level Redo of Image and Video with Adjustable Strength

Diffusion models with prompt and reference image-guided editing have seen rapid progress, yet they remain too coarse for pixel-level control. One promising direction is to incorporate a soft mask that specifies spatially varying edit strengths but training such fine-grained control demands expensive pixel-wise annotati...

Candi Zheng, Yuan Lan · 0 citations
#artificial intelligence Preprint Open access Oct 2026

LEGO-OPD: Factorized Teacher Composition for Multimodal On-Policy Distillation

Multimodal on-policy distillation (OPD) aims to improve visual grounding while preserving the strong reasoning capabilities of language models. Recent multi-teacher approaches combine LLM and VLM teachers to provide complementary supervision. However, directly using a VLM's full predictive distribution entangles its vi...

Jaeyun Shin, Hangeol Chang, Jong Chul Ye · 0 citations
#artificial intelligence Preprint Sep 2026

DriftOPD: Sequence-Level Reverse-KL Distillation for One-Step VLA Policies

Vision-Language-Action (VLA) models increasingly rely on action experts that generate short action chunks under receding-horizon control. While chunk-level training is convenient across robot embodiments, it optimizes local action likelihood without explicitly accounting for long-horizon task success. Sequence-level re...

Youngjun Jun, Kyumin Choi, Young Min Kim et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Beyond Pixel Reconstruction: Retrieval-Guided Glyph-Aware Restoration for Low-Resource Manchu Historical Documents

Historical Manchu documents preserve invaluable linguistic and cultural heritage, yet their digitization is hindered by severe degradations and the scarcity of paired training data. Existing document restoration methods primarily optimize pixel-level reconstruction, which can produce visually plausible results while fa...

Ting Huang, Dongdong Wang, Mingqiu Liang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Decoding the Disaster: Multi-Task Geospatial Reasoning with Vision-Language Models and Crowdsourced Imagery for Disaster Mapping

Crowdsourced imagery provides timely, fine-grained, street-level observations for disaster mapping, complementing conventional remote sensing imagery (RSI) during emergency response. However, such imagery is often unstructured, spatially ambiguous, and lacks reliable geographic metadata, making manual geolocalization a...

Wen-Ping Yin, Fabian Desuer, Zi-Qi Liu et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

LENS-GRF: Permutation-Invariant Lesion Evidence Network with Gated Residual Fusion for Acne Severity Grading and Multi-Rater Clinical Oracle Analysis

Automated acne severity grading requires both whole-face context and fine-grained lesion evidence. We propose LENS-GRF (Lesion Evidence Network with Set-Transformer and Gated Residual Fusion), an interpretable multi-stage framework for four-class acne severity grading. The method combines Adaptive Facial Skin Segmentat...

Muhammad Muhtasim Shahriar, M. F. Mridha · 0 citations

From tech blogs

See all →
Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.