Agentic artificial intelligence is shifting automation from prediction and recommendation toward autonomous execution. This transition creates an economic externality that is not captured by conventional firm-level investment models. A firm receives much of the productivity benefit from delegating decisions to an artif...
Kwan Hong TAN· International Journal of Eco...· 1 citation
Fisher-IRG yields stronger semantic-versus-nuisance predictive selectivity and generally more reproducible subspaces than covariance-based geometry, while recovering systematically distinct local directions.
Many of the qualities that matter most in how people learn and grow, how someone regulates their emotions, reflects on a setback, or stays aware of others during a difficult conversation, are not directly observable. They have to be inferred from how someone speaks, moves, and sounds over time, and they resist the kind...
Siddhant Jain, Dimitra Tsovaltzi· Companion Publication of the...· 0 citations
We present Waypoint 1.5, a real-time diffusion world model for interactive video generation on consumer-grade hardware. Unlike general video diffusion models, interactive world models (iWMs) must respond to dense user controls under strict latency and throughput constraints. Waypoint 1.5 is pre-trained on 100,000 hours...
Segment Anything Model 3 (SAM 3) provides a strong frozen backbone for concept-prompted segmentation, but applying it directly to open-vocabulary semantic segmentation (OVSS) is inefficient: full-resolution decoding is typically run over the entire dataset vocabulary, whereas each image contains only a small active sub...
Parameter-efficient fine-tuning often selects trainable parameters before adaptation using architectural heuristics, without accounting for their varying importance during training. We introduce \textbf{FisherAdapTune}, which progressively selects parameter groups based on temporal drift in their Fisher information. Un...
Ghodsiyeh Rostami, Po-Han Chen, Mahdi S. Hosseini· 0 citations
We present AdvantageFlow, a forward-process reinforcement learning (RL) algorithm for rectified flow models. The algorithm minimizes an advantage-weighted prediction loss, which maximizes reward, regularized by the rollout policy, which convexifies the objective and makes its optimization stable. Our objective can be v...
Branislav Kveton, Anup Rao, Subhojyoti Mukherjee et al.· 0 citations
The AVIS framework leverages autoregressive video diffusion models to restore videos in a streaming manner, naturally eliminating latency bottlenecks and achieving a favorable efficiency-performance trade-off, paving the way toward real-time deployment.
Taesung Kwon, Jonghyun Park, Hyungjin Chung et al.· arXiv.org· 0 citations
Observation-Aligned Mask Priors, a framework that learns the distribution of authentic observation masks and uses it to construct context-query partitions for training from incomplete data, is proposed and demonstrated to be an effective alternative to heuristic masking for learning from incomplete physical observation...
Across camera-space and world-space benchmarks, FactorizedHMR remains competitive with strong baselines, with the clearest gains in occlusion-heavy recovery and drift-sensitive world-space metrics.
Training a diffusion model involves two sources of randomness for each data sample: the timestep and the Gaussian noise realization. The timestep has been studied extensively through scheduling and weighting, whereas the impact of the noise realization at a given timestep is still underexplored. In this work, we examin...
Haokai Zhao, Da Xing, Hanqun Cao et al.· 0 citations
Vision-language models (VLMs) may need to forget visual concepts after deployment because of privacy, copyright, licensing, safety, or policy changes. Conventional machine unlearning modifies model parameters, which may be costly or inaccessible for API-only models. Prompt-based suppression offers a training-free alter...
Zhangyun Tan, Zeliang Zhang, Jiani Liu et al.· 0 citations
Exploring how generative AI could make machine vision more accessible to businesses. The post GenEye in a Box: Making Machine Vision Something You Can Just Ask For appeared first on GPT-Lab.
Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity. The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.
Computer scientist, entrepreneur, and philanthropist will collaborate with the MIT Schwarzman College of Computing to advance AI and scientific discovery.