Skip to content

Category

computer vision

2,913 papers

#artificial intelligence Preprint Open access Oct 2026

Deep Learning for Longitudinal Medical Imaging: A Scoping Review

Longitudinal medical imaging analysis is a cornerstone of modern medical practice and patient care. Deep learning applied to longitudinal imaging offers wide potential to enhance diagnosis and track disease progression by capturing spatial changes over time. With major advances in single-timepoint deep learning for ima...

Francesca Mussa, Divyanshu Tak, Atlas H. Avval et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Autonomous Driving Research Requires a Community-Driven Data Paradigm

Autonomous driving has made remarkable progress, with recent AI advances enabling commercial deployments that are reshaping urban mobility. Yet the field remains far from its universal social promise: autonomous systems that can operate robustly anywhere, anytime, for anyone. We posit that this gap is not merely a mode...

Jinsu Yoo, Zanming Huang, Katie Z Luo et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Pre-training, Reasoning, Benchmarking: X-ray Report Generation on CheXpert Plus Dataset

X-ray image-based Radiology Report Generation (RRG) constitutes a critical research direction within medical artificial intelligence, with great potential to alleviate clinicians' diagnostic workload and shorten patient waiting periods. Despite substantial advances over recent years, the field faces evident bottlenecks...

Xiao Wang, Yuxiang Zhang, Dan Xu et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

GAGR-Lab: Evaluating Joint Spatial-Geometric and Analytic Function Reasoning

Joint spatial-geometric and analytic function reasoning requires translating a perceived spatial configuration into a symbolic function whose executed curve satisfies geometric constraints. We present GAGR-Lab, a framework for measuring this capability through Cartesian game scenes, explicit function semantics, and aut...

Jingyao Zhang, Yun Li, Lu Han · 0 citations
#artificial intelligence Preprint Open access Oct 2026

D3S2: Diffusion-Guided Dataset Distillation for Semantic Segmentation

Dataset distillation (DD) aims to compress large-scale datasets into compact synthetic sets while preserving training efficacy. However, existing studies mainly focus on image classification, leaving dense prediction tasks such as semantic segmentation largely underexplored. In this work, we identify three key challeng...

Wenjie Zheng, Haoji Hu, Jiali Lu et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Adaptive Bidirectional Task Interaction for Joint Segmentation and Classification of Breast Ultrasound

Joint lesion segmentation and tissue classification in breast ultrasound are usually trained with a shared encoder, so the two branches stop exchanging information once their decoders separate. That is exactly where boundary detail and semantic evidence are most complementary. The proposed method restores this exchange...

Abdullah Al Shafi, Md Kawsar Mahmud Khan Zunayed, Safin Ahmmed et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Large Pretraining Datasets Don't Guarantee Robustness after Fine-Tuning in Image Classification

Large-scale pretrained models are widely leveraged as foundations for learning new specialized tasks via fine-tuning, with the goal of maintaining the general performance of the model while allowing it to gain new skills. A valuable goal for all such models is robustness: the ability to perform well on out-of-distribut...

Jaedong Hwang, Brian Cheung, Zhang-Wei Hong et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Ego2World: Compiling Egocentric Cooking Videos into Executable Worlds for Belief-State Planning

Embodied agents in household environments must plan under partial observation: they need to remember objects, track state changes, and recover when actions fail. Existing benchmarks only partially test this ability. Egocentric video datasets capture realistic human activities but remain passive, while interactive simul...

Qinchuan Cheng, Zhantao Gong, Pengzhan Sun et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

4D-HOF: Hand-Object Flow Matching for Feed-Forward 4D Interaction Reconstruction

Existing methods for 4D hand-object reconstruction often rely on costly per-sequence optimization, while generative approaches typically synthesize interactions from random noise, which can lead to unstable interaction prediction. We introduce 4D-HOF, a feed-forward framework that reconstructs 4D hand-object interactio...

Shiqi Li, Sean Cho, Yijie Li et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

DepthWorld: 3D World Model for Robot Manipulation

World models offer a data-driven alternative to traditional simulators for robotics, with applications spanning policy evaluation, improvement, and planning. All of these uses depend on faithful 3D geometry, yet current video-based world models are trained on RGB alone and produce rollouts that look correct frame-by-fr...

Jai Bardhan, Josef Sivic, Vladimir Petrik · 0 citations
#artificial intelligence Preprint Open access Oct 2026

WorldSonus: Bringing Sound to Worlds

Recent advances in world models have enabled increasingly realistic visual synthesis. However, these generated environments remain largely silent. Bringing sound to world models poses three core challenges: real-time generation to keep pace with interactive video streams, interactive control to respond to mid-stream so...

Pengjun Fang, Jingyi Fa, Kam Man Wu et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Selective Transfer of RL Updates for Visual Reasoning

Model merging provides a training-free way to transfer reasoning capabilities from language models to vision-language models (VLMs), but endpoint-based transfer can conflate pre-existing model differences with changes acquired during reasoning post-training. We instead formulate capability transfer around the training-...

Suxin Ji, Hungtao Wan, Mingjun Liu et al. · 0 citations

From tech blogs

See all →
Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.