Skip to content

Category

computer vision

3,022 papers

#computer vision Preprint Aug 2026

TRACE: Training-time Report-guided and Clinically Ordered Concept Editing

Breast ultrasound diagnosis relies on clinically meaningful semantic concepts, yet most deep learning methods adopt end-to-end image-to-label paradigms that lack interpretability and robustness. While concept-based approaches offer a promising alternative, they often assume complete annotations or require multimodal in...

Wen-Tao Yue, Tian-You Lai, Jia-Yu Luo et al. · 0 citations
#computer vision Preprint Open access Oct 2026

Perceptual Color Difference Modeling Using Machine Learning and Human Similarity Judgments

Accurate assessment of color differences is essential for applications ranging from digital design to quality control. While existing color difference metrics, such as CIEDE2000, aim to approximate human perception, they may still exhibit inconsistencies with perceptual judgments. In this study, we investigate a data-d...

Elnara Kadyrgali, Muragul Muratbekova, Adilet Yerkin et al. · 0 citations
#computer vision Preprint Sep 2026

DiFF: Doppler-informed Flow Matching for Human Motion Flow

Perceiving human motion via privacy-preserving 4D millimeter-wave (mmWave) radar is critical for next-generation human-robot interaction (HRI), where point cloud scene flow serves as a foundational motion representation. Yet the extreme sparsity and noise of 4D radar point clouds make non-rigid motion flow estimation s...

Kai Wang, Ming-Le Zhao · 0 citations
#computer vision Preprint Open access Oct 2026

EPIC: Epipolar-Consistent 360{\deg} Immersive Stereo Video Generation

Immersive displays can enable rich and diverse virtual experiences. Manually authoring every possible experience to realize this potential, however, is prohibitively expensive, difficult to scale, and impractical. Generative AI models could remove this bottleneck, but today's models are built for conventional displays...

Debabrata Mandal, Dongdong Fu, Jonathon Miller et al. · 0 citations
#computer vision Preprint Open access Oct 2026

EmAvatar: Multimodal Empathetic Response Generation via Conflict Resolution and Expressive Guidance

Avatar-based multimodal empathetic response generation has emerged as a pivotal capability in human-centric systems, aiming to recognize user emotions and synthesize responses with synchronized text, audio, and talking-face video. Despite recent progress, existing methods still suffer from three critical limitations: (...

Xiaolin Chen, Xuemeng Song, Jinlan Fu et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Structure over Pixels: Learning Variable-Length Visual Programs

Discrete visual tokenizers map images to ordered sequences of tokens, providing a natural representation for structural scene descriptions. Most use a fixed sequence length, while adaptive methods often require post-hoc search or choose among a small set of rates that control the length. We propose STROP, a discrete to...

Piotr Wyrwi\'nski, Kacper Dobek, Krzysztof Krawiec · 0 citations
#machine learning Preprint Open access Oct 2026

Overconfidence and Calibration in Medical VQA: Empirical Findings and Hallucination-Aware Mitigation

As vision-language models (VLMs) are increasingly deployed in clinical decision support, more than accuracy is required: knowing when to trust their predictions is equally critical. Yet, a comprehensive and systematic investigation into the overconfidence of these models remains notably scarce in the medical domain. We...

Ji Young Byun, Young-Jin Park, Jean-Philippe Corbeil et al. · 0 citations
#machine learning Preprint Open access Oct 2026

SOLAR: SVD-Optimized Lifelong Attention for Recommendation

Attention mechanism remains the defining operator in Transformers since it provides expressive global credit assignment, yet its quadratic cost in sequence length N makes long-context modeling expensive and often forces truncation or other heuristics. Linear attention reduces complexity to O(Nd^2) by reordering computa...

Chenghao Zhang, Chao Feng, Yuanhao Pu et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Adaptive Spectral Feature Forecasting for Diffusion Sampling Acceleration

Diffusion models have become the dominant tool for high-fidelity image and video generation, yet are critically bottlenecked by their inference speed due to the numerous iterative passes of Diffusion Transformers. To reduce the exhaustive compute, recent works resort to the feature caching and reusing scheme that skips...

Jiaqi Han, Juntong Shi, Puheng Li et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Two-Step Data Augmentation for Masked Face Detection and Recognition: Turning Fake Masks to Real

The absence of large-scale masked face datasets challenges masked face detection and recognition. We propose a two-step generative data augmentation framework combining rule-based mask warping with unpaired image-to-image translation via GANs, producing masked face samples that go beyond rule-based overlays. Trained on...

Yan Yang, George Bebis, Mircea Nicolescu · 0 citations
#machine learning Preprint Open access Oct 2026

Opportunistic Target Selection: Early Directional Commitment for Query-Efficient Black-Box Adversarial Attacks

Black-box adversarial attacks that minimize only the ground-truth confidence suffer from class drift: perturbations wander through the feature space without committing to a specific adversarial class, wasting queries on diffuse, undirected progress. We introduce Opportunistic Target Selection (OTS), a lightweight wrapp...

Florent Tariolle, Florian Yger · 0 citations
#machine learning Preprint Open access Oct 2026

ReDiF: Resource-Efficient Few-Step Diffusion Distillation via Reinforcement Learning

Step distillation accelerates diffusion sampling by training a few-step student to imitate a many-step teacher, but distillation itself remains expensive. Typically, this requires thousands of GPU-hours and a large pre-generated trajectory dataset. We introduce ReDiF, which casts step distillation as terminal-reward po...

Amirhossein Tighkhorshid, Zahra Dehghanian, Hamid R. Rabiee · 0 citations

From tech blogs

See all →
Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.