Skip to content

Category

machine learning

12,457 papers

#machine learning Preprint Open access Oct 2026

Executing Causal Structure Learning with Linear-Attention Transformers

Transformers can execute algorithms on data given in their input. We ask whether they can do the same for causal discovery. We study a standard continuous method that repeatedly updates a candidate causal graph while enforcing acyclicity. We explicitly construct a fixed-weight transformer whose forward pass exactly rep...

Amartya Roy, Sayar Karmakar · 0 citations
#machine learning Preprint Open access Oct 2026

Kernel Autoresearch for Open-Ended Model Discovery

Kernels encode the inductive bias of a wide range of machine learning models, yet automated kernel design faces a fundamental dilemma. A fixed grammar of base kernels and operators guarantees validity but limits the search to structures expressible by those building blocks. Conversely, unrestricted programs remove this...

Richard Cornelius Suwandi, Feng Yin, Kevin Murphy · 0 citations
#machine learning Preprint Oct 2026

OrBIT: Structure-Guided Embedding Compression

Embedding tables are among the largest components of modern language models. Most compression methods fix a coding geometry such as coordinate blocks, low-rank subspaces, or unrestricted codebooks, and optimize within it. We instead ask whether the coding geometry can itself be discovered. We introduce \emph{OrBIT}, a...

Yunied Puig, Amit Kumar Jaiswal · 0 citations
#machine learning Preprint Open access Oct 2026

Boosting and the Expressive Power of Simple Weak Learners via the $\gamma$-VC Dimension

Boosting converts weak hypotheses with a small edge over random guessing into highly accurate predictors, but the expressive power of the resulting classifier can depend strongly on the structure of the base class. We study this phenomenon through the $\gamma$-VC dimension introduced by Alon et al. (STOC 2021). Our fir...

Arthur da Cunha, Kasper Green Larsen, Liang-Yu Zou · 0 citations
#machine learning Preprint Open access Oct 2026

ResidualQuant: KV Cache Quantization for Looped Transformers with 2-Bit Residuals

Looped Transformers improve parameter efficiency by repeatedly applying shared Transformer blocks over multiple recurrent loops, increasing computational depth without increasing the parameter count. However, KV cache memory still scales with the number of loops, becoming a key memory bottleneck that limits batch size...

Heejun Kim, Junyoung Lee, SangLyul Cho et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Continual Learning without Continual Training

Continual learning requires models to adapt to new domains and new classes while retaining prior knowledge. Many existing methods rely on continued optimization, using regularization, replay, or parameter expansion to prevent new updates from overwriting previously learned knowledge. Instead, we propose replacing conti...

Nikita Narayanan, Ritham Majumdarr, Sonali Parbhoo · 0 citations
#machine learning Preprint Oct 2026

Input-Blind Controls Produce Substantial Oracle Headroom for Layer Programs in Multiple-Choice Evaluation

Adaptive computation aims to improve language-model inference by tailoring execution to each input. For layer programs, oracle evaluations use known answers to estimate the potential gain from this flexibility, before a practical selector is available. However, a gain from selection does not by itself explain why the c...

Yi-Bei Guo, Rui Liu · 0 citations
#machine learning Preprint Open access Oct 2026

Temporally Interpretable Differentiable Decision Trees

Interpretability offers a solution to safe autonomy by providing transparency into an agent's underlying decision-making model. Within sequential-decision making tasks, differentiable decision trees (DDTs) are one approach to such interpretability, maintaining automatic-differentiable policies while providing humans wi...

Eisuke Hirota, Aarav Sane, Rohan Paleja · 0 citations
#machine learning Preprint Open access Oct 2026

Koopman Observers for Diffusion Acceleration: Correcting Feature Forecasts with Shallow Measurements

Feature caching accelerates diffusion sampling by replacing expensive network evaluations with predictions from previously computed activations. However, forecasts based only on past features cannot directly incorporate changes in the current denoising state. We investigate whether inexpensive, freshly computed feature...

Hanru Bai, Yuanchao Xu, Fengyi Li · 0 citations
#machine learning Preprint Oct 2026

ORDERS: An Empirical Study of Norm-Rank Aggregation for Personalized Federated Learning

Personalized federated learning combines shared representations with client-specific predictors, but the contribution of a server weighting rule can be obscured by local training and evaluation choices. We study ORDERS, a configuration that combines a shared backbone, a private residual adapter and classifier, geometri...

Koffka Khan · 0 citations
#machine learning Preprint Open access Oct 2026

AutoAdapt: Automatic Domain Discovery Enables Low-Cost Extensibility

Instruction-tuned models are deployed into environments where domains are heterogeneous and evolve, yet adding new domains or data typically requires costly retraining. We present AutoAdapt, a modular framework that incorporates new domains and data via targeted single-adapter training without modifying other adapters....

Josh McGiff, Salma Mekaoui, Robert Shanahan et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Average-Reward Reinforcement Learning for Multichain MDPs: A Hierarchical Decomposition Approach

We study learning optimal policies in average-reward multichain Markov decision processes (MDPs), where the optimal gain may depend on the initial state and recurrence structures vary across policies, creating challenges for reinforcement learning (RL) methods. We propose an asynchronous value-iteration-based RL algori...

Huizhen Yu, Isaiah Heidt · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Oct 7, 2026

Discovering the value of humanistic inquiry

Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.

Microsoft Research Blog Oct 7, 2026

Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses

Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.