Skip to content

Category

data science

2,518 papers

#machine learning Preprint Sep 2026

Proximal Balancing for Causal Effect Estimation under Unmeasured Confounding

Estimating causal effects from observational data is central to science and policy, but the effects are not identified when confounders are unmeasured. Proximal causal inference addresses this problem with proxies of the unmeasured confounders. However, existing proxy-based approaches either designate proxy roles and s...

Yong-Han Jung · 0 citations
#machine learning Preprint Sep 2026

Amortized Bayesian Inference on Multilevel Models of Arbitrary Structure

We develop a general method for amortized Bayesian inference on multilevel models of arbitrary structure. Given a generative model specified as a directed acyclic graph, our method automatically derives valid factorizations of the joint posterior and matching neural network architectures. The key steps, graph expansion...

Daniel Habermann, Andreas Bulling, Stefan T. Radev et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Estimation of the Label-Noise Transition Matrix with Performance Guarantees via Selective Classification

Modern machine learning depends heavily on massive datasets, but obtaining high-quality annotations at scale is often expensive. As a result, learning from noisily-labeled data has become common, making accurate estimation of the label-noise transition matrix crucial. However, existing transition matrix estimators rely...

Xabier de Juan, Santiago Mazuelas, Yilun Zhu et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Mitigating Representation Gaps in Amortized Bayesian Inference with Auxiliary Supervision

Casting Bayesian inference as a neural network optimization problem targeting an amortized posterior is attractive, as it extends to otherwise intractable statistical models and offers near instantaneous inference for new datasets after prepaying the training cost. Although theory guarantees faithfulness under ideal co...

Hans Olischl\"ager, Svenja Jedhoff, \v{S}imon Kucharsk\'y et al. · 0 citations
#machine learning Preprint Sep 2026

Distributionally robust linear regression through the lens of adversarial training

Distributionally robust optimization (DRO) studies parameter estimation under uncertainty in the underlying probability distribution and has emerged as a principled framework for analyzing robustness and generalization. In particular, Wasserstein DRO, with distributional uncertainty induced by the Wasserstein distance,...

Elis Stefansson, David Vävinggren, Antônio H. Ribeiro · 0 citations
#machine learning Preprint Sep 2026

Principal Component Regression Dominates all Monotone Spectral Filters for Linear Regression

We compare the instance-wise, finite-sample risks of monotone spectral filters for linear regression, a broad class of estimators including principal component regression (PCR), gradient descent (GD), and ridge regression. We show that PCR dominates all monotone spectral filters: compared to any such filter, the risk o...

Juno Kim, Heng-Yu Fu, Peter L. Bartlett et al. · 1 citation
#machine learning Preprint Open access Oct 2026

Towards Optimal Inventory Control under Censored Demand: A Biased Sample-Average Approximation Approach

We study data-driven multi-period lost-sales inventory control under censored demand, where a stockout reveals only that demand exceeded the stocking level. We develop a unified, model-based framework for policy learning from censored data, built on a new cost decomposition for base-stock policies and a biased sample-a...

Yuxuan Han, Xiaoyu Fan, Jiawei Zhang et al. · 0 citations
#machine learning Preprint Sep 2026

A Dynamical Theory of LoRA in Continual Learning

Despite the widespread use of Low-Rank Adaptation (LoRA), little is known about its dynamics in continual learning and the mechanisms by which low-rank updates affect catastrophic forgetting. We provide an asymptotically exact dynamical characterization of LoRA in a solvable two-task teacher-student model. In the high-...

Théo Marchetta, F. Alessandroni, Alessandro Breccia et al. · 0 citations
#machine learning Preprint Sep 2026

Discrete Score Matching Enables Causal Discovery from Count Data

Count data pose a challenge for score-matching-based causal discovery: derivatives are unavailable, and simply replacing them with finite differences does not generally suffice for causal discovery. We generalize SCORE's constant-curvature criterion (Rolland et al., 2022) by conditioning on the node's value, yielding t...

Euijong Song, Hyewon Park, Gunwoong Park · 0 citations
#machine learning Preprint Sep 2026

Minimax Additive Regression under Unknown Dependent Designs

We study additive regression under an unknown and potentially non product design distribution, allowing the number of covariates to grow with the sample size. We consider a coupled class that separately controls the smoothness of the marginal densities and of each additive component multiplied by the corresponding marg...

Baptiste Ferrere, Fabrice Gamboa, Jean-Michel Loubes · 0 citations
#machine learning Preprint Open access Oct 2026

Asymptotic Properties of Support Vector Machines in High-Dimension, Low-Sample-Size Settings under a Spiked Model

In this paper, we consider asymptotic properties of the support vector machine (SVM) in high-dimension, low-sample-size (HDLSS) settings under a spiked model. The existing theory of the SVM in the HDLSS context relies on the geometric representation of HDLSS data, which requires that the eigenvalues of the covariance m...

Yugo Nakayama · 0 citations
#machine learning Preprint Sep 2026

Steepest Guidance: A Practical and Principled Approach to Inference-Time Alignment of Flow and Diffusion-based Models

Inference-time alignment of flow and diffusion-based models is critical for achieving flexible generative modeling. Theoretically, Doob's $h$-transform provides an elegant solution to this problem, and most existing methods are based on this principle. However, in practice, estimating the optimal guidance derived from...

Shokichi Takakura, Akifumi Wachi, Rei Higuchi et al. · 0 citations

From tech blogs

See all →
Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.