Skip to content

Category

machine learning

12,457 papers

#machine learning Preprint Open access Oct 2026

Rounding in Preconditioner Space: Redesigning 4-bit AdamW Optimizer-State Quantization

Quantizing AdamW's optimizer states reduces persistent storage, but quantization errors propagate through the moment recurrences and perturb subsequent adaptive updates. We redesign 4-bit optimizer-state quantization for AdamW from the perspective of \emph{rounding space}: the coordinate in which a quantizer chooses be...

Hanyang Li, Shao Tang, Daniel Thomas Braithwaite et al. · 0 citations
#machine learning Preprint Open access Oct 2026

A Unified Bellman Operator for Safety-Critical Reinforcement Learning

Reinforcement learning in safety-critical domains requires maximizing task performance while strictly adhering to safety constraints. Existing safe reinforcement learning paradigms typically force a trade-off: they either require a priori knowledge to provide strict safety guarantees (e.g., safety filters), or they ena...

Nishanth Arun Rao, Royina Karegoudra Jayanth, Benjamin Eysenbach et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Learning Kilometer-Scale Weather Prediction with Global-Regional Alignment

Kilometer-scale regional weather forecasting is essential for local weather warnings and weather-sensitive decisions. Existing data-driven approaches often rely on numerical forecasts for large-scale guidance or require additional training of global forecasting components. Pretrained global weather models offer an effi...

Guowen Li, Yang Liu, Yujie Wang et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Prospective Prediction of OOD Degradation from Source-Side Training Dynamics

We study whether persistent out-of-distribution (OOD) degradation can be predicted before it is directly observed using only source-side training dynamics. In a controlled shortcut-learning setting, a simple logistic regression predictor develops a clear prospective signal, while training time alone does not. Temporal...

Sasha (Alexander), Monin · 0 citations
#machine learning Preprint Open access Oct 2026

Long Text to Predictive Features: LLM-Guided Blockwise Feature Engineering via Executable Program Search

Industrial risk-control systems typically rely on structured-data models for efficient prediction, yet substantial valuable information remains embedded in unstructured long text. Extracting this information through manual feature engineering is labor-intensive, while requiring a large language model (LLM) to process e...

Ziming Dai, Dabiao Ma, Ziheng Guo et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Marformer: A Transformer for Predicting Missing Data Distributions

Real decisions are made under incomplete information. If we observe only some of the random variables we need, we can predict the others. The \textbf{conditional marginals} over the missing variables are the key ingredient for computing Bayes risk and Value of Information (VOI), the expected gain from acquiring one mor...

Prabhav Singh, Xiheng Tom Wang, Haojun Shi et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Bilevel optimization for data-driven learning of Koopman embeddings using kernel-based autoencoders

Koopman operator theory provides a linear framework for analyzing nonlinear dynamical systems and has become a major tool for data-driven modeling. A central challenge, however, is that finite-dimensional approximations computed by methods such as extended dynamic mode decomposition (EDMD) require the dictionary to be...

Joel-Pascal Ntwali N'konzi, Feliks N\"{u}ske, Stefan Klus · 0 citations
#machine learning Preprint Open access Oct 2026

Closing the Horizon Gap in Policy Optimization for Adversarial MDPs

We consider policy optimization for online episodic tabular Markov decision processes (MDPs) with adversarial losses and bandit feedback. Policy optimization updates the policy locally at each state and avoids optimization over the occupancy-measure polytope, but its existing regret bounds are larger by a factor of the...

Mingyi Li, Taira Tsuchiya · 0 citations
#machine learning Preprint Open access Oct 2026

SplitJEPA: Learning Invariant and Variant Latent Worlds without Reconstruction

Understanding a dynamical world calls for more than a latent state that summarizes its observations: the state should also be organized into the factors that stay shared across related observations and the factors that vary between them. For example, a robot pushing a cube to a goal should take the same action when the...

Ruijin Hua, Zichuan Liu, Zhuokai Zhao et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Ambient Discrete Diffusion: Using the Wrong Data at the Right Time for Data Efficient Learning

We introduce RefineMix, a framework for training discrete diffusion models under severe data scarcity, a common constraint in scientific applications. RefineMix uses out-of-distribution data at selected diffusion times to improve generalization without biasing the sampling distribution. Although this strategy has been...

Julian Kleutgens, Mauricio Tec, Claudio Battiloro et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Composite Online-to-Nonconvex Conversion with Optimal Oracle Complexity

We consider stochastic nonsmooth nonconvex composite optimization, which includes several important problems such as constrained optimization and the regularized training of neural networks. The objective is the sum of a possibly nonsmooth nonconvex Lipschitz function and a convex regularizer, and the function is acces...

Mingyi Li, Taira Tsuchiya, Kenji Yamanishi · 0 citations
#machine learning Preprint Open access Oct 2026

SparseDecoding: Decoding-Aware Pruning for Accurate and Efficient LLM Inference

The memory-bound nature of the decoding stage of large language model (LLM) inference incurs significant latency. Layer-wise training-free network pruning approaches guided by the Hessian have been a prominent solution to this problem, as pruning reduces the number of nonzero parameters read from memory during decoding...

Qitong Wang, Xinwei Niu, Mingluo Su et al. · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Oct 7, 2026

Discovering the value of humanistic inquiry

Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.

Microsoft Research Blog Oct 7, 2026

Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses

Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.