Skip to content

Category

data science

2,429 papers

#machine learning Preprint Open access Oct 2026

Is $\sqrt{d}$ Separation Necessary for Gradient EM to Learn Gaussian Mixtures in High Dimensions?

Learning Gaussian mixture models (GMMs) using the Expectation-Maximization (EM) algorithm and its gradient-based variants is a fundamental problem in machine learning. It is known that randomly initialized (gradient) EM fails to learn multi-component GMMs in the exact-parameterized setting, where the number of componen...

Yiran Zhang, Mo Zhou, Weihang Xu et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Two-Sample Testing for Random Graphs without Vertex Correspondence

Two populations of graphs often have to be compared without any correspondence between their vertices, for instance when networks come from different communities, or when a graph generative model is evaluated against held-out graphs. We study how many graphs such an unaligned two-sample test needs, and which graph stat...

Soham Dan · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Does Muon Need Fine-Grained Spectral Shaping?

Muon combines current and past gradients into matrix momentum. For $M=U\Sigma V^\top$, the idealized polar update $Q=UV^\top$ gives every singular direction the same weight. We refer to this as the flat profile. Several recent optimizers replace this flat profile with fine-grained spectral maps that give each direction...

Meher Chaitanya, Tianyi Zhou, Aristides Gionis · 0 citations
#machine learning Preprint Oct 2026

Bayesian Optimization on Function Spaces via Sparse RKHS Manifolds

Bayesian Optimization (BO) has become an established methodology for minimizing black-box functions of a vector input. Often, however, this parameter vector arises from the discretization of an inherently functional relationship. Several recent articles have considered the Functional Bayesian Optimization (FBO) setting...

Davide Sartor, Meghan E. Huber, Donghyun Kim et al. · 0 citations
#machine learning Preprint Open access Oct 2026

A perspective note on likelihood approximation and inference for complex simulation models using a chain of aggregated normalizing flows

We present a new perspective on the problem of likelihood approximation within the framework of simulation-based inference that promotes scalable and controllable simulation routines for large-scale data analysis, allows efficient parameter space exploration or smooth interpolation in high-dimensions and, thus, support...

Getachew K Befekadu · 0 citations
#artificial intelligence Preprint Open access Oct 2026

DeepAJM: Deep Association Joint Model for Irregularly Sampled data

Joint Models simultaneously model longitudinal and survival outcomes, leveraging patterns in patients' longitudinal trajectory to improve the prediction of survival outcomes. The classical parametric joint models, however, rely on fixed parametric assumptions, making them susceptible to bias under model misspecificatio...

Barsha Halder, Jeffrey A. Thompson · 0 citations
#machine learning Preprint Open access Oct 2026

HyperNSDE: Personalized Neural SDEs for Joint Static-Longitudinal Clinical Data Generation

Synthetic patient data generation is a promising solution to the dual challenge of data scarcity and privacy constraints in healthcare machine learning. Realistic synthesis of patient-level clinical data requires jointly modeling heterogeneous static covariates, irregularly sampled longitudinal trajectories, and inform...

Perrine Chassat, Agathe Guilloux · 0 citations
#machine learning Preprint Oct 2026

Assumption-lean logistic regression with missing covariates

Missing covariates are frequently encountered in supervised learning problems, and classical methods for estimation using such data use carefully chosen imputation schemes for missing data, or likelihood approximations that lead to nonconvex $M$-estimation problems. These methods and their relatives are suitable for sc...

Jyotishka Ray Choudhury, K. A. Verchand, R. Samworth et al. · 0 citations
#machine learning Preprint Oct 2026

How Inefficient Is Natural Gradient Descent? From Exact Optimality to \Theta ( \sqrt{ \log d } ) Divergence

Natural gradient descent (NGD) underlies common methods in ML. For dually flat families, idealized NGD on the forward Kullback--Leibler objective follows the mixture geodesic which is often longer than the shortest Fisher--Rao path. We quantify this overhead by the inefficiency ratio \(R \ge 1\), the Fisher length of t...

Guni Sharon, A. Kuhnle · 0 citations
#machine learning Preprint Open access Oct 2026

A Query Is Not a Commitment: Learning to Correct Expert Answers in Online Deferral

An inaccurate expert can still provide useful information after correction. We study online learning to defer in which the learner chooses an expert and fixes a correction function before purchasing its answer, then applies that function to the answer received. The difficulty is that observed losses reflect both expert...

Yannis Montreuil, Axel Carlier, Lai Xing Ng et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Learning Decision-Stump Thresholds in Context: Dynamics of Softmax Attention

Estimating a decision threshold requires locating observations near an unknown boundary. We study how gradient-based pretraining learns this statistical rule in a two-parameter softmax-attention model with a fixed feature and inequality direction. Pretraining uses labeled contexts and their true thresholds; a fresh thr...

Hong Ha Le, Jackie Lok, Atsushi Nitanda et al. · 0 citations
#machine learning Preprint Open access Oct 2026

sHAIL-Causal: A Sequential Staircase Procedure for Invariant Causal Predictor Discovery

We introduce sHAIL-Causal, the causal specialization of the Saturated Hierarchical Atomic Incremental Learning (sHAIL) paradigm: a sequential staircase procedure that ascends a nested hierarchy of hypothesis classes H_0 < H_1 < ... < H_K once a saturation signal indicates that mastery of the current stage has plateaued...

Ernest Fokou\'e · 0 citations

From tech blogs

See all →
Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.