Skip to content

Category

computer vision

2,913 papers

#artificial intelligence Preprint Oct 2026

MedCORE: Criteria-Grounded Clinical Reasoning for Interpretable Medical Image Diagnosis

Clinical diagnosis is inherently a structured reasoning process, yet existing deep learning models often bypass this structure by mapping image features directly to disease labels without explicitly interrogating the morphological and textural criteria that clinicians systematically evaluate. This limits diagnostic tra...

Asim Khan, Samee Ullah Khan, Dwarikanath Mahapatra · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Knee3DVLM: Dual-Sequence Full-Volume Vision-Language Modeling for Comprehensive Knee MRI Assessment

Vision-language models (VLMs) are increasingly being applied to three-dimensional medical imaging, but their application to knee MRI remains limited, particularly for interpreting the complementary sequences used in clinical practice. We introduce Knee3DVLM, a sequence-aware VLM that uses full-volume DESS and fluid-sen...

Maryam Baizhigitova, Andrew Seohwan Yu, Po-Hao Chen et al. · 0 citations
#artificial intelligence Preprint Oct 2026

GeoPID: Decomposing and Steering Visual Information in Vision-Language Models

While recent vision-language models (VLMs) have shown outstanding performance across diverse applications, they tend to under-use visual information and over-rely on textual context. In this work, we propose \textsc{GeoPID}, a training-free framework that analyzes multimodal information within VLMs from a geometric per...

Seulgi Kim, Zhixiong Zhang, Xin-Wei Zhang et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Transferable Spatial Temporal Coherence Adversarial Attack on Black-Box Vision Language Models for Autonomous Driving

The rapid integration of Vision Language Models (VLMs) into sensitive systems introduces critical safety vulnerabilities that remain unexplored in exist studies. While adversarial attack robustness has been extensively studied for image-based models, the susceptibility of VLMs to temporally-aware adversarial attacks ag...

Heyam Bin Jahlan Areej Alhothali Abeer Alhothali · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Supermarket Product Detection and Recognition: Utilizing Deep Learning with Rectified Imagery

Product Identification has sprung up to become one of the most challenging problems in the automation of the retail industry. With the new industry 5.0 standards, automated inventory management, and catalog creation tasks are vitally important. Object identification models have emerged as a viable answer with their unp...

Mayank Sah, Jimson Mathew · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Decide Before You Look: Learning Which Retrieved Memories Deserve Pixels

Multimodal assistants answer questions from long-term memories that contain images. After retrieval, each retrieved image reaches the answering model either as pixels, at about a thousand visual tokens per image, or as a stored text proxy that often misses the detail the question asks about. We find that the benefit of...

Youxing LI · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Dynamic Alignment and Calibration for Multimodal Learning

Dynamic multimodal learning aims to learn robust representations by adaptively modeling information discrepancies across modalities. However, existing methods still suffer from two limitations: (i) static cross-modal alignment strategies usually impose uniform constraints on all samples while overlooking sample-wise va...

Jinghao Xu, Zhenhua Guo, Xiaofeng Zhu et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Multimodal Knowledge Distillation for Gastric Adenocarcinoma Classification from Whole-Slide Images

Gastric adenocarcinoma (GA) is a leading cause of cancer-related mortality worldwide, and accurate histopathological subtype classification from whole-slide images (WSIs) is essential for effective treatment planning. While multimodal approaches that integrate pathology report text with WSIs can improve classification,...

Shrihari Dumbre, Bikash Santra · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Diverse Motion Customization via Control-based Dynamic Optimization

Despite recent advances in video generation, motion customization remains challenging due to content leakage, where appearance attributes from the reference video unintentionally propagate into the generated output. We identify this issue as a consequence of the generative process collapsing toward the reference video,...

Youngyoon Choi, Kihyun Kim, Jeongwoo Shin et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Later Is Better: Token Reduction for ViTs Under Distribution Shift

Training-free token reduction accelerates vision transformers by removing redundant tokens across layers, recovering most of the original accuracy at a fraction of the compute. These methods, however, are designed and evaluated primarily on clean data, and under real-world distribution shift their accuracy gap to the u...

Hyeongheon Cha, Hyungjun Yoon, Sung-Ju Lee · 0 citations
#artificial intelligence Preprint Open access Oct 2026

What Frame-Level Labels Can and Cannot Do for Small-UAV Point Detection in Thermal Video

The growing use of unmanned aerial vehicles (UAVs) has increased the importance of image-based UAV detection. Learning-based detectors are trained on imagery and annotations, with annotation type determining the information available during training. We focus on learning localization from frame-level target presence/ab...

Wonbin Son, Gyumum Choi, Junil Seo et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Disentangling Dual Image References in Frequency Aware Diffusion Models for Personalized Generation

Personalized image generation aims to synthesize text-driven images conditioned on reference images, while mainly casting the generation as image customization for foreground and style transfer for background. Previous arts of diffusion models suffers from the text misalignment with background for image customization a...

Haipeng Liu, Yang Wang, Meng Wang · 0 citations

From tech blogs

See all →
Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.