This v2 revision reports measured results on the completed implementation of DroneShield-AI, a unified open framework integrating six processing layers: RF signal classification, acoustic motor-signature detection, YOLOv8-based visual detection, evidence-weighted sensor fusion, a Behavioral Intent Classification Engine...
Video retrieval at scale is central to data curation and safety validation in autonomous driving, where users want to find not only scenes but also dynamic events such as cut-ins and hard braking. Existing vision-language and keyword-based retrieval methods often miss these events because the relevant motion may not be...
Manyi Yao, Sparsh Garg, Christian Shelton et al.· 0 citations
Vision-language models (VLMs) achieve high accuracy on many visual question-answering benchmarks, yet it remains unclear whether this accuracy reflects reliable use of the visual details a question depends on. We examine this tension through the lens of visual ignorance: cases where a model can represent task-relevant...
This work studies how moving known computation outside the network changes learning under algebraically equivalent flow interfaces with JLT, a latent Transformer in a frozen variational autoencoder (VAE) representation.
Funing Fu, Tenghui Wang, Guanyu Zhou et al.· 1 citation
Reach audiences
Advertise in front of researchers, engineers, and readers.
Experiments show that COVER improves the coverage-conflict trade-off, making CM-EVS a sparse, compact, and auditable RGB-D-pose resource for geometry-consistent panoramic 3D learning.
While the overall inference latency of Video Diffusion Transformers (DiTs) can be substantially reduced through model distillation, per-step inference latency remains a critical bottleneck. Existing acceleration paradigms primarily exploit redundancy across the denoising trajectory; however, we identify a limitation wh...
Jian Tang, Jiawei Fan, Qingbin Liu et al.· 0 citations
Conformal Prediction (CP) is a distribution-free method for constructing prediction sets with marginal finite-sample coverage guarantees, making it a suitable framework for reliable uncertainty quantification in safety-critical object detection. However, object detection introduces structured multi-output predictions,...
Christoph Ries, Moussa Kassem Sbeyti, N. Bianco et al.· arXiv.org· 0 citations
Many urban prediction tasks depend not only on the static attributes of a location, such as the visual appearance of its built environment, but also on its dynamic function: how people navigate and use it over time. Predicting congestion, travel demand, or road safety risks, for instance, requires mobility information...
Maria Despoina Siampou, Gengchen Mai, Ni Lao et al.· 0 citations
A Confident Learning framework adapted to segmentation for auditing label bias directly in the training data without a clean, unbiased ground truth is presented, and it is shown that label bias influences subgroup separability in the encoder's feature space, an artifact the authors leverage for bias mitigation rather t...
Aditya Parikh, Stella Frank, Snehayan Das et al.· arXiv.org· 2 citations
Post-training quantization (PTQ) is a practical path to deploy large diffusion models, but quantization noise can accumulate over the denoising trajectory and degrade generation quality. We propose Q-Drift, a sampler-side correction that aims to preserve the intended sampling marginals through a deterministic drift adj...
Multimodal instruction tuning is often compute-inefficient because training budgets are spread across large mixed image-video pools whose utility is highly uneven. We present Goal-Driven Data Optimization (GDO), a framework that computes six sample descriptors for each candidate and constructs optimized 1$\times$ train...
Pose-VLA is proposed, a decoupled paradigm that separates VLA training into a pre-training phase for extracting universal 3D spatial priors in a unified camera-centric space, and a post-training phase for efficient embodiment alignment within robot-specific action space.
Exploring how generative AI could make machine vision more accessible to businesses. The post GenEye in a Box: Making Machine Vision Something You Can Just Ask For appeared first on GPT-Lab.
Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity. The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.
Computer scientist, entrepreneur, and philanthropist will collaborate with the MIT Schwarzman College of Computing to advance AI and scientific discovery.