Skip to content

Category

machine learning

12,457 papers

#machine learning Preprint Open access Oct 2026

Measuring and Mitigating Solution Mode Collapse in RLVR

A language model (LM) can usually answer the same question in more than one way, but reinforcement learning with verifiable rewards (RLVR) is indifferent to which correct answer a model produces. A solution will earn the same reward whether it is the thousandth copy of a familiar answer or one the model has never produ...

Liv G. d'Aliberti, Marwa Abdulhai, Sofiia Druchyna et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Low-rank tensor structure of precipitation and its application to satellite-reference merging

The intermittent and variable nature of precipitation makes its accurate estimation over extended domains difficult, yet its spatiotemporal structure suggests that a low-rank representation may be possible. This work represents daily precipitation over the contiguous United States (CONUS) as spatiotemporal tensors and...

Ryan Solgi, Rohan Shankar, Hugo A. Loaiciga · 0 citations
#machine learning Preprint Open access Oct 2026

F$^3$NO: Frequency-Decomposed Finite-Time Flow-map Neural Operators with Cross-Scale Conditioning

Neural operators enable fast PDE forecasting, but repeated predictions accumulate errors and fine-scale structures remain difficult to resolve. We introduce a frequency-decomposed finite-time flow-map neural operator (F$^3$NO) that leverages updated low-frequency features to guide nonlinear refinement of high-frequency...

Fan Wu, Cheng Jing, Kookjin Lee · 0 citations
#machine learning Preprint Open access Oct 2026

Invariant-Measure Reasoners: Stable Representations for Latent Reasoning

Latent reasoning models repeatedly update a latent state using the same recurrent block. As the recurrent depth increases, the sequence of latent states may converge to a compact subset of the state space without necessarily converging to a fixed point. Existing models typically predict by applying a prediction head to...

Yuto Inui, Takuya Konishi, Yoshinobu Kawahara · 0 citations
#machine learning Preprint Open access Oct 2026

Multi-Bandwidth Distribution Matching Distillation: On the Equivalence of Distribution Matching Distillation and Drifting Models

Researchers are exploring effective one-step generative model continuously, and, Drifting Models (Deng et al., 2026), demonstrate great potential in one-step generation recently. There are works that reveal the connection between Diffusion & Flow Style Generative Models (DFSGMs) (Ho et al., 2020; Song et al., 2020a;b;...

Jialin Zhu, Xing Liu, Feixiang He et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Transferability of Learned States in Neural PDE Solvers

Assessing useful reuse in neural PDE solvers is challenging: final accuracy can reflect source learning and target-time computation. Our reuse contract separates solution accuracy, learning contribution, and numerical utility through paired state comparisons, matched target information and budgets, and cost accounting....

Shunye Wang, Haochen Wen, Shuo Li Liu et al. · 0 citations
#machine learning Preprint Open access Oct 2026

ASPIRE: Saddle-Point Discovery through Set Prediction and Physical Refinement

Predicting thermally activated diffusion and defect evolution with event-driven models requires identifying atomic rearrangement mechanisms and their activation barriers. Discovering the associated saddle points is a major computational bottleneck: multiple rearrangements may originate from one metastable state, while...

Yucheng Zhao, Quanyou Zhang, Shaoxiang Qin et al. · 0 citations
#machine learning Preprint Oct 2026

Spectrally Targeted Muon

The Muon optimizer orthogonalizes each update matrix, setting all of its singular values to one, and has proven highly effective for training large language models. It remains unclear, however, whether this success comes from amplifying small singular directions that gradient descent neglects or from suppressing large,...

Vishrut Goyal, Rohan Ramkumar · 0 citations
#machine learning Preprint Open access Oct 2026

SPD-MetaFormer is what you need for small-data brain decoding

Brain signal decoding is challenging because neural recordings are noisy and vary across individuals, while labeled data are often limited. Recent attention-based models on the symmetric positive definite (SPD) manifold have nevertheless achieved strong performance using covariance and connectivity representations, yet...

Zhida Wang, Wei Lyu, Guo Yu et al. · 0 citations
#machine learning Preprint Open access Oct 2026

RH-Detect: A Unified Benchmark for Reward Hacking Detection

Reward hacking, where a model exploits an evaluation signal without completing the intended task, threatens the reliability of deployed language model systems. Existing datasets use different labels, response formats, and metadata conventions, making detector results difficult to compare. We present RH-Detect, a benchm...

Junwei Quan, Evgenii Opryshko, Rohan Subramani et al. · 0 citations
#machine learning Preprint Open access Oct 2026

World-Model Policy Arbiter for Goal-Conditioned Reinforcement Learning

Offline goal-conditioned reinforcement learning (GCRL) has produced a diverse set of goal-reaching algorithms, yet no single algorithm performs best across environments, goals, and even different phases of the same task. Rather than deploying only the best-performing policy, we ask whether a set of frozen goal-conditio...

Junwei Quan, Evgenii Opryshko, Nicholas Rhinehart et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Coefficient Calibration as Selection Pressure in Symbolic Regression

In memetic symbolic regression, candidate structures are compared after coefficient calibration, so the calibration protocol itself contributes to evolutionary selection. Standard centralized calibration evaluates each structure at its pooled-sample optimum, ignoring how stable this calibration is under covariate shift...

Mattia Billa, Veronica Guidetti, Federica Mandreoli · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Oct 7, 2026

Discovering the value of humanistic inquiry

Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.

Microsoft Research Blog Oct 7, 2026

Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses

Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.