Skip to content

Category

data science

2,380 papers

#machine learning Preprint Open access Oct 2026

Expected Sample Complexity in Multi-Armed Bandits

Sample complexity is a widely used metric in sequential decision-making problems, defined as the number of suboptimal decisions during the interaction between the agent and an environment. We study the sample complexity of stochastic multi-armed bandit problems and introduce the expected sample complexity performance m...

Nadav Sukenik, Nadav Merlis · 0 citations
#machine learning Preprint Open access Oct 2026

Eigenvalues of the Hessian in Deep Learning: The Origin of Symmetry and Its Breaking

Hessian spectra at trained models in deep learning exhibit a persistent pattern: eigenvalues organize into distinct clusters, including a large bulk near zero and a few isolated outliers. This paper shows that a natural account of these spectral phenomena emerges when the original setting is understood as a departure f...

Yossi Arjevani · 0 citations
#machine learning Preprint Open access Oct 2026

Identifiability of a dissipative knowledge-dynamics model: exact recovery under designed excitation, degeneration on observational data

Human learning is a dissipative dynamical process: mastery accumulates through practice, decays through forgetting, and propagates across interdependent concepts. We model it as a nonlinear dissipative system of ordinary differential equations whose parameters are mechanistically meaningful (a concept-transfer matrix e...

Arman Kostanian, Armen Beklaryan · 0 citations
#machine learning Preprint Open access Oct 2026

AdaPS-LiNGAM: Adaptive Predecessor Selection for Linear Non-Gaussian Acyclic Models under Small-Sample Settings

Causal discovery becomes particularly challenging when the available sample size is small relative to the number of variables. This challenge also arises in the linear non-Gaussian acyclic model (LiNGAM), an identifiable framework for causal discovery from observational data. DirectLiNGAM estimates a causal order, whic...

Shun Yanashima, Kentaro Kanamori, Hirofumi Suzuki · 0 citations
#machine learning Preprint Open access Oct 2026

Fluctuations of Nonlinear Observables in Mean Field Neural Network Training

Mean field limits describe the training dynamics of wide neural networks through the evolution of the empirical distribution of their parameters. Although functional central limit theorems characterize the asymptotic fluctuations of this distribution, quantities of practical interest are typically nonlinear observables...

Arnaud Descours (UCBL), Geoffrey Lacour (MaIAGE) · 0 citations
#machine learning Preprint Open access Oct 2026

Leaner Transformers Can Easily Learn to Cluster

Transformers have in-context learning capabilities, where some known learning algorithms can be executed in the forward pass through the model. Recent work shows that transformers can exactly perform Lloyd's algorithm for $k$-means clustering with $n$ points in $d$ dimensions with an embedding size $d_{\textsf{emb}} =...

Charlotte Park, Kenneth L. Clarkson, Lior Horesh et al. · 0 citations
#machine learning Preprint Open access Oct 2026

EntroPrefill: Renyi-Guided Context Pruning with Conditional Stability Guarantees for Retrieval-Augmented Generation

Mid-prefill pruning can reduce the sequence processed by deeper transformer layers, but attention concentration alone does not certify that discarded context is dispensable. We formulate EntroPrefill as a Renyi-guided proposal mechanism coupled to explicit constraints on discarded attention mass. Sink-isolated, regular...

Inbasekaran S · 0 citations
#machine learning Preprint Open access Oct 2026

Gauss-Newton Accuracy and Indefinite Hessians: Uniform Coexistence in Low-Cost Sets

We study the accuracy of Gauss-Newton curvature in ridge-regularized nonlinear least squares. Under local regularity and persistence of level-set curvature magnitude along an exact-fit section, we prove uniform coexistence of two curvature regimes. Global minimizers exist, and every global minimizer has relative Hessia...

Kihun Rhee, Hanjoon Byun · 0 citations
#machine learning Preprint Open access Oct 2026

Global Exponential Convergence of Two-Layer Linear Network Training

We prove global exponential (linear) convergence with an explicit rate in the rich scaling for wide two-layer linear networks trained with smooth Polyak-Lojasiewicz predictor losses. Gradient flow in the factors closes exactly in terms of a finite-dimensional Bures flow of the neuron law covariance, in which the predic...

Stephen Y Zhang, Gabriel Peyr\'e · 0 citations
#machine learning Preprint Open access Oct 2026

Benign Overfitting under Heterogeneous Input Fusion

Benign overfitting is extensively studied when learning from a single high-dimensional input, but its behavior under heterogeneous input fusion remains largely unexplored. We study this question for minimum-norm linear interpolation under a heterogeneous Gaussian design, comparing two statistically dependent input bloc...

Houzhen Liu, Xiaobo Xia · 0 citations
#machine learning Preprint Open access Oct 2026

The Symbol of the Surrogate: Measuring Numerical Provenance in Neural PDE Solvers

Neural PDE surrogates are trained on numerical solver outputs that contain both physical evolution and solver-specific discretization errors. Because surrogates are also evaluated against held-out trajectories from the same solver, standard benchmarks cannot distinguish fidelity to the exact evolution from imitation of...

Ridham Patel · 0 citations
#machine learning Preprint Open access Oct 2026

An Accuracy--Information Tradeoff for Loss-Difference Conditional Mutual Information

Loss-difference conditional mutual information (ld-CMI) uses the smallest of the standard observations in the supersample hierarchy of generalization bounds: it measures what a learner's loss differences reveal about which candidate of each pair it was trained on. Accuracy is known to force information into the model;...

Hazar Yueksel · 0 citations

From tech blogs

See all →
Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.