Skip to content

Category

machine learning

12,457 papers

#artificial intelligence Preprint Oct 2026

Convex-Concave Reinforcement Learning

Policy learning drives many of the most consequential and heavily-invested applications of reinforcement learning today. Yet the core optimization problem it rests on (maximizing expected return) is notoriously non-convex, even under a direct policy parameterization, and the field has largely responded by avoiding it:...

Shripad Deshmukh, Yaswanth Chittepu, Dhawal Gupta et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

MaRK: Markov-adapted Recurrent Kernels for Dynamic Operator Conditioning in State Space Models

State Space Models (SSMs) offer an efficient alternative to Transformers for sequence modeling, yet conditioning pre-trained SSMs for iterative generation typically operates outside the recurrent operator, through input injection or activation modulation. While such mechanisms expose the model to conditioning informati...

Syed Ibrahim Omer, Ginny Y. Wong. Xiangyu Zhao · 0 citations
#artificial intelligence Preprint Oct 2026

U-Space: Uncovering When and Why Uncertainty Arises in Language Models

Large language models are informing decisions with ever-higher stakes. As the consequences of their errors grow, a central question becomes harder to ignore: how much can we trust an individual answer? Yet recognizing when to defer remains difficult because language models can present incorrect conclusions with fluent...

Tobias Braun, Nils Loose, Alexander Herzog et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Talking with Language Models

When we interact with large language models (LLMs), are we having a conversation? They are designed to invite us to treat them as intelligent interlocutors who remember, act, and make commitments. But appearances deceive. We introduce the artifactual stance, a framework that reconceives human-AI interaction as artifact...

James Ravi Kirkpatrick, Alexandru Radulescu, Rachel Katharine Sterken · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Careful Judge: Safe and Efficient Human-AI Collaborative Decision Making

In human-AI collaborative decision making, human review can prevent unsafe AI decisions, but each human judgment is costly. Treating human intervention after AI abstention as a one-off fallback misses the opportunity to improve future AI decisions for greater automation, yet AI adaptively learning from selectively quer...

Chenyu Zhang, Rachel Luo, Boyi Li et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Quad-State Safety Evaluation of Open-Weight Large Language Models on Non-Canonical Inputs

Standard safety evaluations of large language models assess harmful requests written in canonical plain text, while models in real-world deployment routinely receive inputs containing emojis, altered spellings, encoded strings, and character-level variations. This work introduces the Adversarial Surface-Form Robustness...

Pavan Maddula · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Algorithmic Scratchpads and Curriculum Staging for Arithmetic Reasoning in Tiny Transformers

Autoregressive Large Language Models (LLMs) frequently struggle with deterministic multi-step algorithmic tasks such as multi-digit multiplication and long division. In this paper, we investigate the mechanics of multi-step arithmetic in compact "Tiny" Transformers (~10.6M non-embedding parameters, 49.3M total) trained...

Sourabh Kasliwal · 0 citations
#artificial intelligence Preprint Open access Oct 2026

LASER: Latent Space Adjoint Matching for Support-Constrained Entropy-Regularized Offline RL

While offline reinforcement learning (RL) enables policy optimization from static datasets without costly online interaction, it remains bottlenecked by the risk of executing out-of-distribution (OOD) actions. Recent approaches mitigate this by learning a behavior-cloning policy through flow matching and then performin...

Songyuan Zhang, Oswin So, Eric Yang Yu et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Work While They Sleep: Exploiting Evaluation Latency for Fully Bayesian Optimization

Black-box optimization problems are ubiquitous across science and engineering, often dealing with expensive objective functions. This objective latency has two consequences during optimization: (i) the objective evaluation dominates execution time, and (ii) sample-efficient algorithms are crucial to accelerate developm...

Gustavo Sutter, Alejandro Comas-Leon, David Holzm\"uller et al. · 0 citations
#artificial intelligence Preprint Oct 2026

On KL-Regularized Policy Optimization

Asynchronous reinforcement learning (RL) for large language model (LLM) agents trains one policy on trajectories generated by another: rollouts come from stale checkpoints, and the inference engine's probabilities differ from the trainer's even at identical parameters. Standard remedies either clip importance ratios, w...

Yi-Fan Zhang · 0 citations
#artificial intelligence Preprint Open access Oct 2026

GraphOPD: Graph-Augmented On-Policy Distillation for LLM Agents

On-policy distillation post-trains large language model agents by supplying dense, step-level guidance from a teacher policy when the reinforcement-learning reward is sparse and arrives only once per trajectory. Existing instantiations allocate this guidance by the size of the teacher-student divergence at each step, o...

Bohan Lin, Liyi Chen, Zhuoning Guo et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

FinVector-Market-4B: A Controlled Study of LoRA Adaptation for Structured Financial Tasks

FinVector-Market-4B adapts Qwen/Qwen3.5-4B with rank-16 LoRA on a 22,000-example corpus for structured financial tasks. We evaluate the base and adapted models on the same 600-example benchmark under implicit and explicit JSON-schema contracts. Supplying the schema alone raises base-model JSON validity from 0% to 91.3%...

Alina Khaybullina · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Oct 7, 2026

Discovering the value of humanistic inquiry

Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.

Microsoft Research Blog Oct 7, 2026

Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses

Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.