Skip to content

Category

computer vision

3,022 papers

DroneShield-AI: A Multi-Modal Sensor Fusion Framework for Real-Time Autonomous Drone Threat Detection, Behavioral Intent Classification, and Swarm Intelligence in Contested Airspace

This v2 revision reports measured results on the completed implementation of DroneShield-AI, a unified open framework integrating six processing layers: RF signal classification, acoustic motor-signature detection, YOLOv8-based visual detection, evidence-weighted sensor fusion, a Behavioral Intent Classification Engine...

Marius Bayizere · 0 citations
#machine learning Preprint Open access Sep 2026

Driving Video Retrieval for Complex Queries with Structured Grounding

Video retrieval at scale is central to data curation and safety validation in autonomous driving, where users want to find not only scenes but also dynamic events such as cut-ins and hard braking. Existing vision-language and keyword-based retrieval methods often miss these events because the relevant motion may not be...

Manyi Yao, Sparsh Garg, Christian Shelton et al. · 0 citations
#machine learning Preprint Open access Sep 2026

Diagnosing Visual Ignorance in Vision-Language Models

Vision-language models (VLMs) achieve high accuracy on many visual question-answering benchmarks, yet it remains unclear whether this accuracy reflects reliable use of the visual details a question depends on. We examine this tension through the lens of visual ignorance: cases where a model can represent task-relevant...

Runyu Zhou, Qi Zhang, Yisen Wang · 0 citations
#machine learning Preprint Open access Sep 2026

FIS-DiT: Breaking the Few-Step Video Inference Barrier via Training-Free Frame Interleaved Sparsity

While the overall inference latency of Video Diffusion Transformers (DiTs) can be substantially reduced through model distillation, per-step inference latency remains a critical bottleneck. Existing acceleration paradigms primarily exploit redundancy across the denoising trajectory; however, we identify a limitation wh...

Jian Tang, Jiawei Fan, Qingbin Liu et al. · 0 citations

Probabilistic Object Detection with Conformal Prediction

Conformal Prediction (CP) is a distribution-free method for constructing prediction sets with marginal finite-sample coverage guarantees, making it a suitable framework for reliable uncertainty quantification in safety-critical object detection. However, object detection introduces structured multi-output predictions,...

Christoph Ries, Moussa Kassem Sbeyti, N. Bianco et al. · 0 citations
#machine learning Preprint Open access Sep 2026

TrajGANR: Trajectory-Centric Urban Multimodal Learning via Geospatially Aligned Neural Representations

Many urban prediction tasks depend not only on the static attributes of a location, such as the visual appearance of its built environment, but also on its dynamic function: how people navigate and use it over time. Predicting congestion, travel demand, or road safety risks, for instance, requires mobility information...

Maria Despoina Siampou, Gengchen Mai, Ni Lao et al. · 0 citations

Towards Fairness under Label Bias in Image Segmentation: Impact, Measurement and Mitigation

A Confident Learning framework adapted to segmentation for auditing label bias directly in the training data without a clean, unbiased ground truth is presented, and it is shown that label bias influences subgroup separability in the encoder's feature space, an artifact the authors leverage for bias mitigation rather t...

Aditya Parikh, Stella Frank, Snehayan Das et al. · 2 citations
#machine learning Preprint Open access Sep 2026

Q-Drift: Quantization-Aware Drift Correction for Diffusion Model Sampling

Post-training quantization (PTQ) is a practical path to deploy large diffusion models, but quantization noise can accumulate over the denoising trajectory and degrade generation quality. We propose Q-Drift, a sampler-side correction that aims to preserve the intended sampling marginals through a deterministic drift adj...

Sooyoung Ryu, Mathieu Salzmann, Saqib Javed · 0 citations
#machine learning Preprint Open access Sep 2026

Less Data, Faster Convergence: Goal-Driven Data Optimization for Multimodal Instruction Tuning

Multimodal instruction tuning is often compute-inefficient because training budgets are spread across large mixed image-video pools whose utility is highly uneven. We present Goal-Driven Data Optimization (GDO), a framework that computes six sample descriptors for each candidate and constructs optimized 1$\times$ train...

Rujie Wu, Haozhe Zhao, Hai Ci et al. · 0 citations

Universal Pose Pretraining for Generalizable Vision-Language-Action Policies

Pose-VLA is proposed, a decoupled paradigm that separates VLA training into a pre-training phase for extracting universal 3D spatial priors in a unified camera-centric space, and a post-training phase for efficient embodiment alignment within robot-specific action space.

Haitao Lin, Hanyang Yu, Jingshun Huang et al. · 11 citations

From tech blogs

See all →
Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.