Skip to content

Category

machine learning

12,457 papers

#artificial intelligence Preprint Open access Oct 2026

A Positive Case for Faithfulness: LLM Self-Explanations Help Predict Model Behavior

LLM self-explanations are often presented as a promising tool for AI oversight, yet their faithfulness to the model's true reasoning process is poorly understood. Existing faithfulness metrics have critical limitations, typically relying on identifying unfaithfulness via adversarial prompting or detecting reasoning err...

Harry Mayne, Justin Singh Kang, Dewi Gould et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Requirement-Based Testing: Enhancing Reinforcement Learning with Game Theory

We consider the automatic online synthesis of black-box test cases from functional requirements specified as automata for reactive implementations. The goal of the tester is to reach some given state, so as to satisfy a coverage criterion, while monitoring the violation of the requirements. We develop an approach based...

Ocan Sankur (DEVINE), Thierry J\'eron (DEVINE), Nicolas Markey (DEVINE) et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Decoupling Exploration from Optimization in RLVR

Modern language models undergo reinforcement learning with verifiable rewards (RLVR) on top of already-trained checkpoints. A key promise of RLVR is the discovery of new reasoning strategies. In principle, a model can sample novel ideas absent from its prior training data. In practice, however, augmenting RLVR with str...

Saif Punjwani, Micah Goldblum · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Composing What Each Teacher Learned: Multi-Teacher On-Policy Distillation through Teacher-Relative Shifts

Multi-teacher on-policy distillation (MOPD) is used in two settings. In common-domain composition, several teachers score each student rollout from one prompt domain and their signals form a single target; in routed-domain distillation, prompts from different domains are assigned to the corresponding specialist. Both s...

Hejian Sang, Zhengze Zhou, Shayan Mohajer Hamidi et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

A Good Self-Teacher Meets the Student Where They Are: Joint On-Policy Learning and Teaching

Reinforcement Learning (RL) from outcome rewards suffers from sparse supervision, particularly on difficult, long-horizon tasks where successful trajectories are rare and costly to generate. On-Policy Distillation (OPD) offers an attractive alternative by providing dense token-level supervision from a stronger teacher...

Randy Ardywibowo, Arnav Dalal, Jiantao Jiao · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Q-Learning with Scalar Adjoint Matching

Flow policies capture rich and diverse action distributions, and fine-tuning them with off-policy RL to improve beyond the demonstrations has drawn growing interest. However, fine-tuning a flow policy against a learned value function is not trivial, because the policy generates its action over many flow steps. Adjoint...

Yonghoon Dong, Minsung Yoon, Jaehyuk Kim et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Training Parallel Speculative Draft Models by Directly Minimizing Expected Decoding Rounds

Speculative decoding accelerates large language model inference by using a low-cost draft model to propose tokens that the full-size target model verifies in parallel. Parallel and semi-autoregressive (semi- AR) drafters improve drafting efficiency by proposing an entire block in a single forward pass, but training the...

Yunxiao Zhao, Changxiao Cai · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Fault-tolerant foundation models

Emerging computer hardware often trades reliability for energy efficiency; here we show that large-language models (LLMs) can be trained to tolerate this unreliability, and that rather than degrading, their error resilience actually increases as they grow. Modified neural scaling laws inferred from 40,000 GPU-hours of...

Trevor McCourt, Ila R. Fiete, Isaac L. Chuang · 0 citations
#artificial intelligence Preprint Open access Oct 2026

SemanticFold: Latent Sequence Compression SeparatesLanguage Modeling, Decodability, and Reasoning

We study whether latent sequence compression of prompt prefixes preserves the capabilities that large language models rely on during inference. We introduce SemanticFold, a compression scheme that folds prefix hidden states at learned boundaries, and evaluate it across five model scales: Qwen3-1.7B, Qwen3-8B, SmolLM2-1...

Mingyan Liu, Min Huang · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Logarithmic Regret via Passive Change Detection in Piecewise-Stationary Self-Tuning Regulation

We study minimum-variance control of an unknown autoregressive system with exogenous inputs and coefficients that change at unknown times. Under bounded independent disturbances, fixed detection gaps, stability and feasibility conditions, and sufficient time between changes, we prove \(O((C+1)\log((T+1)/\delta))\) regr...

A. Ch. Madhusudanarao, Rahul Singh · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Stationary Bias and Extrapolation in Nonlinear Two-Timescale Stochastic Approximation

Constant-step stochastic approximation generally has a nonzero stationary mean error that persists under time averaging. This paper studies that error for nonlinear two-timescale recursions driven by an exogenous finite-state Markov chain. Under stated smoothness assumptions and conditions on the stationary distributio...

A. Ch. Madhusudanarao, Rahul Singh · 0 citations
#artificial intelligence Preprint Open access Oct 2026

From Prompts to Trees: Effective LLM-Guided Tree Generation for Few-Shot Tabular Classification

While Large Language Models (LLMs) possess rich world knowledge and impressive generalization capabilities, their direct application to tabular data classification is hindered by high inference costs and limited interpretability. In contrast, decision trees are fast and transparent but often underperform in low-data re...

Yue Qiu, Zekang Du, Yiqun Diao et al. · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Oct 7, 2026

Discovering the value of humanistic inquiry

Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.

Microsoft Research Blog Oct 7, 2026

Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses

Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.