Skip to content

Category

data science

2,380 papers

#machine learning Preprint Open access Oct 2026

Quadratic Weak-to-Strong Generalization in Random Feature Networks via Random Matrix Theory

Weak-to-strong generalization is the phenomenon where a strong student model trained with labels produced by a weak teacher model is able to generalize better than the teacher. In this paper, we study this phenomenon in two-layer random feature networks where the model strength is determined by its width. Using tools f...

Deborah Oliveira, Elliot Paquette · 0 citations
#machine learning Preprint Open access Oct 2026

Trust-Region Optimization for Smooth Potential-Interaction Energies in Wasserstein Space

Finding low-energy configurations of interacting particles and approximating probability distributions lead to the minimization of potential-interaction energies in Wasserstein space. These energies can be nonconvex, making it important to exploit second-order information while controlling the reliability of local appr...

You Wan, Ting Gao, Jinqiao Duan · 0 citations
#machine learning Preprint Open access Oct 2026

Slow Beats Fast at the Kesten-Stigum Threshold: Minimax, Fisher-Information and Belief-Propagation Characterizations of the Information-Computation Gap in Sparse Stochastic Block Models

We study community recovery in the sparse symmetric stochastic block model with $q$ communities, average degree $d$ and signal strength $\lambda$ through statistical decision theory and Fisher information, and obtain three characterizations of the Kesten-Stigum threshold $d\lambda^2=1$ and of the information-computatio...

Soroor Ghandali · 0 citations
#machine learning Preprint Open access Oct 2026

Why Forget-Only Unlearning Needs Memorization

Machine unlearning asks for a deletion algorithm whose output is close to retraining from scratch without the selected forget examples. In this work, we study forget-only unlearning, where the deletion algorithm receives only the trained model and the examples to forget, with no retained data or extra training informat...

Luka Radi\'c, Vikrant Singhal, Amartya Sanyal · 0 citations
#machine learning Preprint Open access Oct 2026

Oracle-Efficient and Parameter-Free Agnostic Smoothed Online Learning

Online learning is an attractive framework in many domains because it permits well-defined learning even when data are dependent or chosen adversarially. This generality, however, comes at a steep price, introducing significant statistical and computational barriers. Recently, smoothed online learning has emerged as a...

Sasha Voitovych, Adam Block, Alexander Rakhlin et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Two-Level Softmax Sampling Done Right: Correcting Bias from Size Imbalance and Dispersion

Sampling from a softmax distribution is a fundamental operation in machine learning, but its linear complexity in the number of items makes exact sampling impractical at scale. Two-level softmax (2LS) sampling is a popular alternative enabling sublinear-time sampling. Assuming items are partitioned into clusters, 2LS f...

Walid Bendada, Guillaume Salha-Galvan · 0 citations
#machine learning Preprint Open access Oct 2026

Koopman Observers for Diffusion Acceleration: Correcting Feature Forecasts with Shallow Measurements

Feature caching accelerates diffusion sampling by replacing expensive network evaluations with predictions from previously computed activations. However, forecasts based only on past features cannot directly incorporate changes in the current denoising state. We investigate whether inexpensive, freshly computed feature...

Hanru Bai, Yuanchao Xu, Fengyi Li · 0 citations
#machine learning Preprint Oct 2026

Revisiting Explainable AI through Model-Independent Concept Dictionaries

Modern applications of AI rely on increasingly complex models. Explainable AI (XAI) has emerged as a set of techniques aimed at improving model transparency. However, existing XAI methods typically assume input features to be inherently interpretable, or they rely on intermediate internal abstractions that are difficul...

T. Schnake, Doreen Schöppenthau, Alexander Meyer et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Shared Gaussianization: What Gaussian Regularizers Certify About Contrastive Learning, and What They Miss

What can a distribution-matching regularizer such as SIGReg in LeJEPA certify about contrastive learning? We study shared Gaussianization (SG), a characteristic-function Gaussianity test on the average of two normalized views, scaled by an independent $\chi_d$ radius. Because disagreeing views shorten the average, one...

Ruoyu Zhao, Yuting Chen, Jinheng Zhang et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Finite-Sample Approximation of Hessian-Guided Perturbed Wasserstein Gradient Flows

Wasserstein gradient flow extends gradient descent to probability measures. Its Hessian-guided perturbed variant (PWGF) adds Gaussian perturbations to escape saddle points in nonconvex problems. We investigate when its approximation by finitely many interacting particles remains accurate over growing time horizons. Our...

Ryotaro Kawata, Atsushi Nitanda, Taiji Suzuki · 0 citations
#machine learning Preprint Open access Oct 2026

Pre-training of Bayesian Optimization Algorithm through Bayesian Optimization

Bayesian optimization (BO) is widely used as a standard approach for expensive black-box optimization. However, BO algorithms often involve parameters that must be specified in advance, and their performance can strongly depend on these choices. We propose a framework for optimizing such parameters using sample paths d...

Satoshi Katayama, Shoyo Hunt, Shintaro Masuda et al. · 0 citations
#machine learning Preprint Open access Oct 2026

m-Set Adversarial Bandits with Winner Feedback

We show upper and lower bounds on the regret of $m$-set adversarial bandits for different utilities (winner reward or sum of rewards) and feedback models (winner index, winner reward, sum of rewards, and their combinations). By comparing to standard bounds for combinatorial and MNL bandits, our results reveal how subtl...

Nicol\`o Cesa-Bianchi, Matteo Papini · 0 citations

From tech blogs

See all →
Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.