Skip to content

Category

data science

2,429 papers

#machine learning Preprint Open access Oct 2026

Anytime-valid simulation-based hypothesis testing

For a given data distribution $(X_t)_{t \in \mathbb{N}} \sim Q$ i.i.d., we investigate the hypothesis testing problem: $H_0: Q = P_0$ vs. $H_1: Q = P_1$, for two different model probability distributions $P_0$ and $P_1$. In contrast to the standard setting, where analytic densities $p_0$ and $p_1$ are given, here, we c...

Patrick Forr\'e, Lydia Brenner · 0 citations
#machine learning Preprint Open access Oct 2026

Beyond Marginal Monitoring: Distributed Joint-Distribution Testing for Data Concept Drift in Large Scale E-Commerce Operations

Concept drift threatens production machine learning, yet the empirical behavior of multivariate two-sample drift detectors at scale remains under-characterized. Existing benchmarks rarely address the hundreds of millions of rows and high-cardinality features typical of industrial-operational datasets. We evaluate five...

Cagdas Pullu, Mahmut Emir Arslan, Bugra Balkac et al. · 0 citations
#machine learning Preprint Open access Oct 2026

ProximalFM: Amortized Proximal Causal Inference under Hidden Confounding

Standard causal identification methods often assume no unmeasured confounding and can fail when relevant confounders are unobserved. Proximal causal inference instead uses proxy variables to identify effects under hidden confounding. However, nonparametric proximal estimation can be challenging in practice: recovering...

Christophe Muller, Ayub Kharel, Alex Luedtke et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Stochastic Gradient Descent Ascent is Suboptimal for Nonconvex-PL Min-Max Games

How far can stochastic gradient descent ascent (SGDA) go by tuning its timescale ratio and step sizes in nonconvex min-max games? We answer this question for nonconvex-PL (NC-PL) games by establishing the first tight complexity of two-timescale SGDA with a fixed timescale ratio and non-increasing step sizes. For $\ell$...

Junsoo Ha · 0 citations
#machine learning Preprint Oct 2026

Trustworthy Method Comparison with AI Judges: Estimation and Design under Order, Batch, and Aggregation Effects

Large language models (LLMs) are increasingly used as judges for automated AI evaluation. A common practice is to randomize prompt sequences and average the resulting scores, but its statistical validity remains unclear. We show that LLM evaluation mechanisms can be approximated by a class of Markov generalized linear...

Tian-Xi Li, Jie Ding · 0 citations
#machine learning Preprint Open access Oct 2026

High-dimensional online calibration from harmonic weights

We study the online calibration of multidimensional forecasts over an arbitrary convex set $Y\subseteq\mathbb{R}^d$ relative to an arbitrary error norm $\|\cdot\|_{L}$. For forecasting $d$ binary outcomes simultaneously ($Y=[0,1]^d$), we give the first algorithm that achieves $\varepsilon$-calibration in a number of ro...

Maxwell Fishelson, Mehryar Mohri · 0 citations
#machine learning Preprint Open access Oct 2026

Nash Social Welfare for Multi Armed Bandits: Trajectory-wise Expected and High Probability Regret

We study fair multi-armed bandits under the Nash Social Welfare (NSW) objective, which measures performance via the geometric mean of accumulated rewards. Existing work defines Nash regret as $\mathrm{NR}_T = \mu^\star - (\prod_{t=1}^T \mathbb{E}\mu_{I_t})^{1/T}$, where $\mu_{I_t}$ is the mean reward of the recommended...

Avishek Ghosh · 0 citations
#machine learning Preprint Open access Oct 2026

Exact Calibration and Sharp Risk Geometry for Volume-Sampled Ridge Regression

We study ridge regression from exactly $s$ distinct rows of a fixed design. Responses are fixed, and only the subset is random. The determinant law and selected ridge fit share one positive definite penalty. Established mean identities and exponential-family duality give the unique penalty that matches a prescribed ful...

Kihun Rhee · 0 citations
#machine learning Preprint Open access Oct 2026

Stability of Measure-to-Measure Transformers on Sub-Gaussian Data

Transformers have exhibited impressive empirical success across various domains, but their theoretical foundations remain less developed. This work constitutes a mathematical study of the measure-to-measure operators defined by transformers. We show that transformers map sub-Gaussian inputs to sub-Gaussian outputs; thi...

Frank Cole, Nicholas H. Nelsen, Takashi Furuya · 0 citations
#machine learning Preprint Open access Oct 2026

Uniform Discrete Diffusion Models are Minimax Optimal for Estimating Distributions with Small Effective Support Size

Discrete diffusion models have emerged as a practically successful framework for generative modeling on discrete product spaces, yet their statistical generalization properties remain poorly understood. Discrete real-world data such as text or biological sequences often concentrate on a small fraction of the astronomic...

Dongsun Yoon, Saptarshi Chakraborty · 0 citations
#machine learning Preprint Open access Oct 2026

Asymptotic Analysis of Empirical Risk Minimization on Entry-wise i.i.d. Heavy-Tailed Data

Many real-world datasets exhibit unusually large values far more frequently than predicted by Gaussian models. Heavy-tailed distributions capture this behavior, yet evaluating learning performance under them remains challenging because rare, large feature entries retain non-vanishing effects even in high dimensions. Ev...

Kaito Takanami, Takashi Takahashi, Yoshiyuki Kabashima · 0 citations
#machine learning Preprint Open access Oct 2026

Explicit Asymptotic Bounds for Sequential Calibration Beyond $T^{2/3}$

Probability forecasts are calibrated when predicted probabilities match empirical outcome frequencies: among events assigned a probability $p$, we'd hope that the fraction of positive outcomes is close to $p$. We study the problem of sequential forecasting of binary outcomes. The classical $O(T^{2/3})$ bound on expecte...

Eric Dai, Maxwell Fishelson · 0 citations

From tech blogs

See all →
Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.