Skip to content

Category

machine learning

12,457 papers

#machine learning Preprint Oct 2026

Expected Sample Complexity in Multi-Armed Bandits

Sample complexity is a widely used metric in sequential decision-making problems, defined as the number of suboptimal decisions during the interaction between the agent and an environment. We study the sample complexity of stochastic multi-armed bandit problems and introduce the expected sample complexity performance m...

Nadav Sukenik, Nadav Merlis · 0 citations
#machine learning Preprint Open access Oct 2026

Eigenvalues of the Hessian in Deep Learning: The Origin of Symmetry and Its Breaking

Hessian spectra at trained models in deep learning exhibit a persistent pattern: eigenvalues organize into distinct clusters, including a large bulk near zero and a few isolated outliers. This paper shows that a natural account of these spectral phenomena emerges when the original setting is understood as a departure f...

Yossi Arjevani · 0 citations
#machine learning Preprint Open access Oct 2026

NeuralZip: Reusable Setup for Fast Lossless Compression

Lossless compression can reduce the storage and movement of model weights without changing their floating-point values, but repeated statistical analysis and code construction add computational overhead. We study whether the statistical structure of exponents can be prepared once and reused. For this, we introduce Neur...

Mart\'in Bravo, Samuel Horv\'ath, Gonzalo Navarro et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Identifiability of a dissipative knowledge-dynamics model: exact recovery under designed excitation, degeneration on observational data

Human learning is a dissipative dynamical process: mastery accumulates through practice, decays through forgetting, and propagates across interdependent concepts. We model it as a nonlinear dissipative system of ordinary differential equations whose parameters are mechanistically meaningful (a concept-transfer matrix e...

Arman Kostanian, Armen Beklaryan · 0 citations
#machine learning Preprint Open access Oct 2026

Layerwise Error Attribution for Fast and Robust Mixed-Precision Post-Training Quantization

Mixed-precision post-training quantization is a network compression method that assigns bits layer by layer, under a global memory budget using a small calibration set. The main difficulties are to overcome the combinatorial nature of the allocation problem and to manage the sensitivity to small, potentially corrupted...

Samy Houache (IMB, UB), Yann Traonmilin (IMB et al. · 0 citations
#machine learning Preprint Oct 2026

MUNITE: Unified Multimodal Latent Inference for Any-to-Any Multimodal Generation

We introduce MUNITE, a latent-variable framework for flexible any-to-any multimodal generation that treats encoding and latent generation as the same inference problem under different amounts of observed evidence. Given any subset of modalities, MUNITE models the conditional distribution over the latent representation...

Kyeongmin Yeo, Minhyuk Sung · 0 citations
#machine learning Preprint Open access Oct 2026

For Those Who Believe in Faithfulness: Optimizing the Area Under Insertion and Deletion Curves for Ranking Relative Feature Importance

The adoption of machine learning for socially relevant tasks requires effective explainable artificial intelligence (XAI) methods to better understand the behavior of machine learning models. Attribution methods are a popular XAI approach in which input-output relationships are characterized by heat maps that reflect t...

Bj{\o}rn Leth M{\o}ller, Bulat Ibragimov, Christian Igel · 0 citations
#machine learning Preprint Open access Oct 2026

Dual-QK: Sharp Queries and Flat Keys for Prunable 2-bit KV Caches

Long inputs and extended generation increase the storage and access costs of the key-value (KV) cache. Low-bit quantization reduces storage and memory traffic, while query-channel pruning can further reduce key-cache reads. Rotation-based quantization redistributes the energy of key outliers across channels. To maintai...

Sunjoo Whang, Jungjun Oh, Minsung Kim et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Homogenization in Multi-Agent Systems

Multi-agent systems (MAS) leverage interactions between agents to perform complex tasks. Despite their success, we show that these interactions can also lead to homogenization, i.e., agents converging to similar behaviors. Homogenization in MAS can reduce agent diversity and reinforce shared failures. In this paper, we...

Prakhar Ganesh, Kyra Wilson, Luca Zappella et al. · 0 citations
#machine learning Preprint Open access Oct 2026

BoT-GRPO: Efficient Process-Reward RL for Reasoning via Bag-of-Token Aggregation

Reinforcement learning is now central to eliciting reasoning in large language models, while in the popular algorithm Group Relative Policy Optimization (GRPO) every token in a rollout receives the same advantage. We ask how to make process supervision efficient: accelerating convergence and improving final quality wit...

Yingxiang Yang, Weihang Xiao, Zhunxuan Wang et al. · 0 citations
#machine learning Preprint Open access Oct 2026

A Proof-of-Concept Study of Weakly Supervised Labeling of Fine-Grained EEG Components for Artifact Attenuation

Electroencephalography (EEG) is highly susceptible to electromyographic (EMG) artifacts, whose temporal heterogeneity and spatial-spectral overlap with neural activity can leave mixed sources after blind source separation. Existing artifact-removal methods are further limited by scarce reliable component-level ground t...

Lu Wang-N\"oth, Hai Huang, Philipp Heiler et al. · 0 citations
#machine learning Preprint Open access Oct 2026

AdaPS-LiNGAM: Adaptive Predecessor Selection for Linear Non-Gaussian Acyclic Models under Small-Sample Settings

Causal discovery becomes particularly challenging when the available sample size is small relative to the number of variables. This challenge also arises in the linear non-Gaussian acyclic model (LiNGAM), an identifiable framework for causal discovery from observational data. DirectLiNGAM estimates a causal order, whic...

Shun Yanashima, Kentaro Kanamori, Hirofumi Suzuki · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Oct 7, 2026

Discovering the value of humanistic inquiry

Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.

Microsoft Research Blog Oct 7, 2026

Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses

Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.