Skip to content

Category

computer vision

3,022 papers

#artificial intelligence Open access Aug 2026

The Autonomy Externality: A Welfare-Economic Model of Agentic Artificial Intelligence, Correlated Failure, and Optimal Assurance Policy

Agentic artificial intelligence is shifting automation from prediction and recommendation toward autonomous execution. This transition creates an economic externality that is not captured by conventional firm-level investment models. A firm receives much of the productivity benefit from delegating decisions to an artif...

Kwan Hong TAN · 1 citation
#machine learning Book Open access Sep 2026

Multimodal Detection of Higher-Order Behavioral Constructs: Self-Compassion in Structured Reflective Interactions

Many of the qualities that matter most in how people learn and grow, how someone regulates their emotions, reflects on a setback, or stays aware of others during a difficult conversation, are not directly observable. They have to be inferred from how someone speaks, moves, and sounds over time, and they resist the kind...

Siddhant Jain, Dimitra Tsovaltzi · 0 citations
#computer vision Preprint Sep 2026

Waypoint-1.5: A Real-Time Video World Model for Consumer Hardware

We present Waypoint 1.5, a real-time diffusion world model for interactive video generation on consumer-grade hardware. Unlike general video diffusion models, interactive world models (iWMs) must respond to dense user controls under strict latency and throughput constraints. Waypoint 1.5 is pre-trained on 100,000 hours...

Rajit Rajpal, Shahbuland Matiana, Liew Wei Pyn et al. · 0 citations
#artificial intelligence Preprint Open access Sep 2026

ActiveSAM: Fast and Accurate Open-Vocabulary Semantic Segmentation with Frozen SAM 3

Segment Anything Model 3 (SAM 3) provides a strong frozen backbone for concept-prompted segmentation, but applying it directly to open-vocabulary semantic segmentation (OVSS) is inefficient: full-resolution decoding is typically run over the entire dataset vocabulary, whereas each image contains only a small active sub...

Tran Dinh Tien, Zhiqiang Shen · 0 citations
#artificial intelligence Preprint Open access Sep 2026

Fisher-Guided Progressive Parameter Selection for Adaptive Fine-Tuning

Parameter-efficient fine-tuning often selects trainable parameters before adaptation using architectural heuristics, without accounting for their varying importance during training. We introduce \textbf{FisherAdapTune}, which progressively selects parameter groups based on temporal drift in their Fisher information. Un...

Ghodsiyeh Rostami, Po-Han Chen, Mahdi S. Hosseini · 0 citations
#artificial intelligence Preprint Open access Sep 2026

AdvantageFlow: Regularized Advantage-Weighted RL in Flow Models

We present AdvantageFlow, a forward-process reinforcement learning (RL) algorithm for rectified flow models. The algorithm minimizes an advantage-weighted prediction loss, which maximizes reward, regularized by the rollout policy, which convexifies the objective and makes its optimization stable. Our objective can be v...

Branislav Kveton, Anup Rao, Subhojyoti Mukherjee et al. · 0 citations

Accelerating Video Inverse Problem Solvers with Autoregressive Diffusion Models

The AVIS framework leverages autoregressive video diffusion models to restore videos in a streaming manner, naturally eliminating latency bottlenecks and achieving a favorable efficiency-performance trade-off, paving the way toward real-time deployment.

Taesung Kwon, Jonghyun Park, Hyungjin Chung et al. · 0 citations
#artificial intelligence Preprint May 2026

Observation-Aligned Mask Priors for Learning Physical Fields from Authentic Occlusions

Observation-Aligned Mask Priors, a framework that learns the distribution of authentic observation masks and uses it to construct context-query partitions for training from incomplete data, is proposed and demonstrated to be an effective alternative to heuristic masking for learning from incomplete physical observation...

Chi-Yuan Ma, Zi-Han Zhou, Tian Yu · 0 citations
#artificial intelligence Preprint Open access Sep 2026

NoiseRater: Meta-Learned Noise Valuation for Diffusion Model Training

Training a diffusion model involves two sources of randomness for each data sample: the timestep and the Gaussian noise realization. The timestep has been studied extensively through scheduling and weighting, whereas the impact of the noise realization at a given timestep is still underexplored. In this work, we examin...

Haokai Zhao, Da Xing, Hanqun Cao et al. · 0 citations
#artificial intelligence Preprint Open access Sep 2026

Suppression Is Not Forgetting: Residual Recoverability in Visual Concept Unlearning for VLMs

Vision-language models (VLMs) may need to forget visual concepts after deployment because of privacy, copyright, licensing, safety, or policy changes. Conventional machine unlearning modifies model parameters, which may be costly or inaccessible for API-only models. Prompt-based suppression offers a training-free alter...

Zhangyun Tan, Zeliang Zhang, Jiani Liu et al. · 0 citations

From tech blogs

See all →
Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.