Skip to content

Category

machine learning

12,457 papers

#machine learning Preprint Open access Oct 2026

Fluctuations of Nonlinear Observables in Mean Field Neural Network Training

Mean field limits describe the training dynamics of wide neural networks through the evolution of the empirical distribution of their parameters. Although functional central limit theorems characterize the asymptotic fluctuations of this distribution, quantities of practical interest are typically nonlinear observables...

Arnaud Descours (UCBL), Geoffrey Lacour (MaIAGE) · 0 citations
#machine learning Preprint Open access Oct 2026

Leaner Transformers Can Easily Learn to Cluster

Transformers have in-context learning capabilities, where some known learning algorithms can be executed in the forward pass through the model. Recent work shows that transformers can exactly perform Lloyd's algorithm for $k$-means clustering with $n$ points in $d$ dimensions with an embedding size $d_{\textsf{emb}} =...

Charlotte Park, Kenneth L. Clarkson, Lior Horesh et al. · 0 citations
#machine learning Preprint Open access Oct 2026

EntroPrefill: Renyi-Guided Context Pruning with Conditional Stability Guarantees for Retrieval-Augmented Generation

Mid-prefill pruning can reduce the sequence processed by deeper transformer layers, but attention concentration alone does not certify that discarded context is dispensable. We formulate EntroPrefill as a Renyi-guided proposal mechanism coupled to explicit constraints on discarded attention mass. Sink-isolated, regular...

Inbasekaran S · 0 citations
#machine learning Preprint Open access Oct 2026

SoftSEEPS improves ML-based precipitation forecasting

In this paper we have developed a differentiable approximation of the well-known SEEPS score, which we name SoftSEEPS. This allows the training of a Machine Learning model to forecast precipitation directly. We test SoftSEEPS on the IMERG dataset (0.1 degree resolution) by training a decoder for precipitation on the la...

Jost Arndt, Utku Isil, Noelia Otero et al. · 0 citations
#machine learning Preprint Open access Oct 2026

EC-EarthFlow: Probabilistic emulation of daily transient global climate model simulations with flow matching

We introduce EC-EarthFlow, a generative flow matching model that emulates simulations from the physical climate model EC-Earth3. The model is trained on transient simulations from EC-Earth3 (1950-2166, SSP2-4.5) to predict the day ahead temperature field from the previous days temperature as well as annual mean tempera...

Kirien Whan, Nikolaj T. M\"ucke, Karin van der Wiel · 0 citations
#machine learning Preprint Open access Oct 2026

Pretraining Shapes Spectral Structure: Architecture- and Strategy-Conditional Prediction of OOD Robustness in Foundation Models

Can we determine whether a foundation model will generalize out-of-distribution (OOD) before any target data is available? Existing diagnostics require source or target data, which rules them out before a target domain exists. Those that use the weights alone apply one statistic to every architecture, and do not separa...

Sangyoon Bae, Sk Miraj Ahmed, Shinjae Yoo et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Gauss-Newton Accuracy and Indefinite Hessians: Uniform Coexistence in Low-Cost Sets

We study the accuracy of Gauss-Newton curvature in ridge-regularized nonlinear least squares. Under local regularity and persistence of level-set curvature magnitude along an exact-fit section, we prove uniform coexistence of two curvature regimes. Global minimizers exist, and every global minimizer has relative Hessia...

Kihun Rhee, Hanjoon Byun · 0 citations
#machine learning Preprint Oct 2026

DSTNet: Dynamic Spectral Trajectory Network for Causal Multi-Horizon Financial Forecasting

Wavelet-based financial forecasters typically use the transform only to denoise, or reduce it to a single spectral snapshot at the forecast origin, and the convolution that produces the coefficients is usually bilateral, so it can read past the forecast origin. DSTNet instead retains the recent evolution of filter-bank...

Aashish Bohra, Lokendra Vishwakarm · 0 citations
#machine learning Preprint Open access Oct 2026

Closed-Form Noise Calibration Against Membership Inference for Random-Allocation DP-SGD

DP-SGD protects training data by adding Gaussian noise to clipped gradients. The amount of noise is usually chosen by running a numerical privacy accountant inside a search. We study DP-SGD with random allocation, where each epoch uses every record once, at a randomly chosen step. For this setting we give a one-line fo...

Murat Bilgehan Ertan, Marten van Dijk · 0 citations
#machine learning Preprint Open access Oct 2026

When Rank Rises as LLMs Degrade

Post-training adapts language models in non-stationary environments. Practitioners monitor representation health with RankMe and related spectral statistics, often assuming that rank falls when representations degrade. We show that this assumption is unsafe for LLM post-training. In a controlled study of Qwen3-0.6B wit...

Zhaohui Geoffrey Wang · 0 citations
#machine learning Preprint Open access Oct 2026

When does a network's training history predict its future learning better than its current state? Evidence from a response probe and a forecasting screen

Networks that behave alike now can still learn differently when training continues. Work on loss of plasticity and critical periods shows that the path to a state shapes what follows; it does not show whether the path carries information that a measurement of the state itself misses. We ask when the training history of...

Martin Hofmann, Patrick M\"ader · 0 citations
#machine learning Preprint Oct 2026

The Identifiability and Observability of Deep Normalized Attention

We study which parameters of deep, unmasked, single-head attention are determined by its input--output function. For known positive nonconstant real-analytic normalizers, the function generically determines the effective scores and combined value map up to the signs induced by even normalizers. This proves the real-ana...

Pranav Venkata Konda · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Oct 7, 2026

Discovering the value of humanistic inquiry

Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.

Microsoft Research Blog Oct 7, 2026

Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses

Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.