Skip to content

Category

machine learning

12,457 papers

#artificial intelligence Preprint Open access Oct 2026

RouterInterp: Understanding Superposed Specialisation in Mixture of Experts Routing

Sparse Mixture of Experts (MoE) models scale more efficiently than dense models by routing tokens to modular expert networks that are only active for processing a fraction of tokens. A leading hypothesis for the performance of MoE models is that each expert specialises in a single, coherent domain. However, interpretab...

Ilya Lasy, Nora Yinuo Cai, Kola Ayonrinde · 0 citations
#artificial intelligence Preprint Oct 2026

Internalizer: Portable Context-to-Parameter Mapping for Very Large Language Models

Hypernetworks that map a context directly to a LoRA adapter let a large language model carry that context in its weights, but prior work has demonstrated them only on base models of up to 14 billion parameters. We present the Internalizer, a state-of-the-art, portable Context-to-Parameter Mapping hypernetwork that gene...

Peter Devine, Nick Ryan, Benjamin Sirb et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Harness Evolution Hits a Ceiling: When Weight Training Should Begin

Improving a long-horizon LLM agent means evolving the harness around a frozen model or training its weights. We let a self-evolving harness make the system stronger first, then cross seed and evolved harnesses with base and trained weights to learn which gains the trained model keeps and which still need the runtime. W...

Yuan Tian, Bing Hu, Hao Wang et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Who Verifies the Verifier? Co-Evolving Inspectable Graders with Self-Improving Agents

We changed the agent: did it actually get better? Every self-improving agent loop answers this hundreds of times, and every answer comes from a verifier. On open-ended tasks none exists, so the loop is handed a hand-written rubric or a bare LLM judge grading output from a model like itself, inviting reward hacking and...

Xing Zhang, Guanghui Wang, Yanwei Cui et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Estimating great expectations under autoregressive language models with potentials

Many applications of language models hinge not on individual samples but on the expectation of a test functional under the model. Estimating such expectations reliably can be computationally expensive. In this paper, we show how to make estimation more efficient by exploiting the next-token conditional probabilities wh...

Francesco I. Re, Shubhangi Ghosh, Tim Vieira et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

RaReCache: Bridging the Gap in Cross-Model KV Cache Reuse via Rank disagreement-based Selective Recomputation

Cross-model KV-cache reuse remains a key challenge in modern LLM serving. Coding agents and multi-model systems increasingly route a shared context across models: a user may switch models mid-session, or a cascade may escalate a difficult query. Because KV caches contain model-specific representations, each switch typi...

Sreetama Sarkar, Saptarshi Mitra, Sitao Huang et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

RL-ARC: Calibrating Large Reasoning Models via Reasoning-guided Uncertainty

Language models (LMs) are commonly trained with Reinforcement Learning with Verifiable Rewards (RLVR) to enhance their reasoning capabilities. However, since RLVR does not explicitly account for calibration during training, it can lead to severe calibration degradation, including overconfidence. Recent calibration-awar...

Gukhyeon Lee, SangKeun Lee · 0 citations
#artificial intelligence Preprint Open access Oct 2026

DivMoE: Fine-Grained MoE Upcycling via Cross-Domain Expert Composition

Mixture-of-Experts (MoE) architectures have become essential for scaling large language models, with recent work demonstrating the benefits of fine-grained expert designs. Training such models from scratch is expensive, and sparse upcycling from pre-trained dense models is an attractive alternative. However, we identif...

Yuxuan Lou, Kai Yang, Geng Zhang et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

When Lower Reconstruction Loss Hurts: Distributionally Robust Refinement for Low-Bit LLM Quantization

Weight-only post-training quantization (PTQ) relies heavily on reconstruction loss minimization to preserve model quality at low precision. We show that the weights favored by minimizing this loss need not yield better model performance on new tasks. In fact, we find that lower reconstruction loss can even degrade mode...

Yanlong Zhao, Xiaoyuan Cheng, Huihang Liu et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

AgentHorizon: Evaluating Agentic Judges for Long-Horizon Computer-Use Tasks

Computer-use agents are capable of completing complex tasks, increasing the use of automatic judges to determine success, either for training or for evaluation without human involvement. Despite their flexibility, their reliability on long tasks spanning multiple applications remains unclear. A trajectory, composed of...

Xing Han L\`u, Dheeraj Vattikonda, Sina Hajimiri et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Curating Always-Loaded Context for LLM Agents: A Capacitated Assortment Model with Censored Feedback

At the start of every session, LLM agents load a fixed context file, such as $\texttt{AGENTS.md}$. Each loaded token in the file is charged again in every later round of the session, and these files can degrade performance as they grow in size. However, in practice, human or automated curators usually grow these files...

Zexuan Liu, Yuning Yang, Tiancheng Zhao · 0 citations
#artificial intelligence Preprint Open access Oct 2026

StoreBench: A Live-Commerce Environment for Evaluating and Training Autonomous Operator Agents

Reinforcement learning environments are now a primary lever for improving large language model (LLM) capabilities in post-training, yet most agentic benchmarks remain static: the world moves only when the agent acts, the reward is a terminal verdict, and the pass bar is set arbitrarily. We introduce StoreBench, a live-...

Daksh Raghuvanshi, Ved Vedere, Yifan Wang · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Oct 7, 2026

Discovering the value of humanistic inquiry

Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.

Microsoft Research Blog Oct 7, 2026

Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses

Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.