Skip to content

Category

data science

2,473 papers

#machine learning Preprint Open access Oct 2026

Zero-Flow Two-Sample Tests

Motivated by the success of modern flow-based generative models in modeling complex data, we study two-sample testing through the lens of flow-based methods. We propose the Zero-Flow Two-Sample Test (ZF2ST), built on the zero-flow criterion, which characterizes distributional equality through a time-reversal antisymmet...

Yakun Wang, Leyang Wang, Song Liu et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Tighter Regret Bounds for Contextual Action-Set Reinforcement Learning

We study episodic reinforcement learning with fixed reward and transition functions, but with episode-dependent admissible action sets that are observed at the start of each episode. Performance is measured by cumulative regret against the episode-wise optimal value, $\sum_{k=1}^K [V^{*,M^k} - V^{\pi^k,M^k}]$, where $M...

Zijun Chen, Zihan Zhang · 0 citations
#machine learning Preprint Open access Oct 2026

Differentiable Expectation-Maximisation and Applications to Gaussian Mixture Model Optimal Transport

The Expectation-Maximisation (EM) algorithm is a central tool in statistics and machine learning, widely used for latent-variable models such as Gaussian Mixture Models (GMMs). Despite its ubiquity, EM is typically treated as a non-differentiable black box, preventing its integration into modern learning pipelines wher...

Samuel Bo\"it\'e, Eloi Tanguy, Julie Delon et al. · 0 citations
#machine learning Preprint Sep 2026

Proximal Balancing for Causal Effect Estimation under Unmeasured Confounding

Estimating causal effects from observational data is central to science and policy, but the effects are not identified when confounders are unmeasured. Proximal causal inference addresses this problem with proxies of the unmeasured confounders. However, existing proxy-based approaches either designate proxy roles and s...

Yong-Han Jung · 0 citations
#machine learning Preprint Sep 2026

Amortized Bayesian Inference on Multilevel Models of Arbitrary Structure

We develop a general method for amortized Bayesian inference on multilevel models of arbitrary structure. Given a generative model specified as a directed acyclic graph, our method automatically derives valid factorizations of the joint posterior and matching neural network architectures. The key steps, graph expansion...

Daniel Habermann, Andreas Bulling, Stefan T. Radev et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Estimation of the Label-Noise Transition Matrix with Performance Guarantees via Selective Classification

Modern machine learning depends heavily on massive datasets, but obtaining high-quality annotations at scale is often expensive. As a result, learning from noisily-labeled data has become common, making accurate estimation of the label-noise transition matrix crucial. However, existing transition matrix estimators rely...

Xabier de Juan, Santiago Mazuelas, Yilun Zhu et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Mitigating Representation Gaps in Amortized Bayesian Inference with Auxiliary Supervision

Casting Bayesian inference as a neural network optimization problem targeting an amortized posterior is attractive, as it extends to otherwise intractable statistical models and offers near instantaneous inference for new datasets after prepaying the training cost. Although theory guarantees faithfulness under ideal co...

Hans Olischl\"ager, Svenja Jedhoff, \v{S}imon Kucharsk\'y et al. · 0 citations
#machine learning Preprint Sep 2026

Distributionally robust linear regression through the lens of adversarial training

Distributionally robust optimization (DRO) studies parameter estimation under uncertainty in the underlying probability distribution and has emerged as a principled framework for analyzing robustness and generalization. In particular, Wasserstein DRO, with distributional uncertainty induced by the Wasserstein distance,...

Elis Stefansson, David Vävinggren, Antônio H. Ribeiro · 0 citations
#machine learning Preprint Sep 2026

Principal Component Regression Dominates all Monotone Spectral Filters for Linear Regression

We compare the instance-wise, finite-sample risks of monotone spectral filters for linear regression, a broad class of estimators including principal component regression (PCR), gradient descent (GD), and ridge regression. We show that PCR dominates all monotone spectral filters: compared to any such filter, the risk o...

Juno Kim, Heng-Yu Fu, Peter L. Bartlett et al. · 1 citation
#machine learning Preprint Open access Oct 2026

Towards Optimal Inventory Control under Censored Demand: A Biased Sample-Average Approximation Approach

We study data-driven multi-period lost-sales inventory control under censored demand, where a stockout reveals only that demand exceeded the stocking level. We develop a unified, model-based framework for policy learning from censored data, built on a new cost decomposition for base-stock policies and a biased sample-a...

Yuxuan Han, Xiaoyu Fan, Jiawei Zhang et al. · 0 citations
#machine learning Preprint Sep 2026

A Dynamical Theory of LoRA in Continual Learning

Despite the widespread use of Low-Rank Adaptation (LoRA), little is known about its dynamics in continual learning and the mechanisms by which low-rank updates affect catastrophic forgetting. We provide an asymptotically exact dynamical characterization of LoRA in a solvable two-task teacher-student model. In the high-...

Théo Marchetta, F. Alessandroni, Alessandro Breccia et al. · 0 citations
#machine learning Preprint Sep 2026

Discrete Score Matching Enables Causal Discovery from Count Data

Count data pose a challenge for score-matching-based causal discovery: derivatives are unavailable, and simply replacing them with finite differences does not generally suffice for causal discovery. We generalize SCORE's constant-curvature criterion (Rolland et al., 2022) by conditioning on the node's value, yielding t...

Euijong Song, Hyewon Park, Gunwoong Park · 0 citations

From tech blogs

See all →
Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.