Skip to content

Category

data science

2,473 papers

#machine learning Preprint Open access Oct 2026

The Benchmarking Epistemology: Validity Theory for Evaluating Machine Learning Models

Predictive benchmarking, evaluating machine learning models based on predictive performance and competitive ranking, is central to machine learning research and scientific inquiry. However, benchmark scores at best measure performance relative to a specific dataset and learning problem. Drawing substantial scientific i...

Timo Freiesleben, Sebastian Zezulka · 0 citations
#machine learning Preprint Open access Oct 2026

Meta-reinforcement learning with minimum attention

Minimum attention applies the least action principle in changes of control concerning state and time, first proposed by Brockett. The involved regularization is highly relevant in emulating biological control, such as motor learning. We apply minimum attention in reinforcement learning (RL) as part of the rewards and i...

Shashank Gupta, Pilhwa Lee · 0 citations
#machine learning Preprint Open access Oct 2026

Stochastic Optimal Control for Continuous-Time fMRI Representation Learning

Learning robust representations from functional magnetic resonance imaging (fMRI) is fundamentally challenged by the temporal irregularity and noise inherent in data from heterogeneous sources. Existing self-supervised learning (SSL) methods often discard critical temporal information by discretizing or averaging fMRI...

Joonhyeong Park, Byoungwoo Park, Chang-Bae Bang et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Roto-translated Local Coordinate Frames For Interacting Dynamical Systems

Modelling interactions is critical in learning complex dynamical systems, namely systems of interacting objects with highly non-linear and time-dependent behaviour. A large class of such systems can be formalized as $\textit{geometric graphs}$, $\textit{i.e.}$, graphs with nodes positioned in the Euclidean space given...

Miltiadis Kofinas, Naveen Shankar Nagaraja, Efstratios Gavves · 0 citations
#machine learning Preprint Oct 2026

Sample complexity bounds for categorical Markov random fields via Discrete Diffusions

Many applications in statistics, economics, and physics require sampling from high-dimensional categorical distributions with local dependence structures. Examples include finite memory language models, Ising and Potts systems in statistical physics and protein folding, etc. In modern machine learning, discrete diffusi...

Shivam Kumar, Nabarun Deb · 0 citations
#machine learning Preprint Oct 2026

Wasserstein Gradient Flows and Forward-Only Diffusion Are Not Enough for Multimodal Sampling

There has been a proliferation of sampling algorithms based on Wasserstein gradient flows (WGF) and forward-only diffusion processes (FODP), often accompanied by theoretical guarantees of exponentially fast convergence to the target distribution. These guarantees are frequently interpreted as evidence that such methods...

Daniel McBride, Pratik Khandagale, Cristina Garcia-Cardona et al. · 0 citations
#machine learning Preprint Oct 2026

Error-Corrected Inference-Time Scaling for Imperfect Diffusion Models

Inference-time scaling adapts pretrained diffusion models to new sampling tasks without additional training. Existing methods rely primarily on Monte Carlo sampling with more particles, yet are premised on the pretrained model being exact. In practice, data and training limitations make the model imperfect, and these m...

Zuo-Kai Wen, Louis Grenioux, E. Weinan et al. · 0 citations
#machine learning Preprint Oct 2026

Generalized Engression Models

We consider estimating the conditional distribution of a multivariate outcome given covariates when its coordinates may be continuous, binary, categorical, ordinal or rankings, and are conditionally dependent on one another. Different statistical methods have been developed for each outcome type, and most of them targe...

Xin-Wei Shen, Zi-Jian Guo, Francis Bach · 0 citations
#machine learning Preprint Oct 2026

Lower Bounds for Stochastic First-Order Algorithms with Variance Reduction in Nonconvex--Concave Minimax Optimization

We establish complexity lower bounds for stochastic first-order algorithms in nonconvex--concave minimax optimization, allowing algorithms to use variance reduction. Our main contribution is a lower bound for a zero-respecting algorithm class that permits variance reduction, extending beyond the algorithmic restriction...

Jia-Yi Song, Zi Xu · 0 citations
#machine learning Preprint Oct 2026

The hidden advantage of mask resampling: a theory of masked autoencoders

Why can masked prediction learn useful representations that unmasked reconstruction misses? We study this question in a high-dimensional model of a masked autoencoder (MAE) trained on data with shared latent structure and heterogeneous noise. We prove that masked linear reconstruction can recover the latent feature at...

Jorge Medina Moreira, Lorenzo Bardone, Lenka Zdeborová · 0 citations
#machine learning Preprint Open access Oct 2026

Zero Flux: Flow-Based Comparison of High-Dimensional Discrete Distributions

Comparing two high-dimensional discrete distributions has always been a challenging task due to the exponentially growing state space and complex changes in interactions. A recent work suggests comparing distributions through a vector field trained using flow matching between two continuous distributions. The resulting...

Leyang Wang, Yakun Wang, Song Liu et al. · 0 citations
#machine learning Preprint Oct 2026

Tolerance-Based Fairness Auditing: Violation Certification and Sensitivity Screening

A unified tolerance-based fairness auditing framework for two complementary auditing objectives: violation certification, which prioritizes control of false violation declarations, and sensitivity screening, which prioritizes reducing missed violations is developed.

Jie Tang, Chuan-Long Xie, Li-Xing Zhu · 0 citations

From tech blogs

See all →
Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.