Skip to content

Category

data science

2,518 papers

#machine learning Preprint Open access Oct 2026

Minimax rates for learning spectral Barron functions by deep ReLU neural networks

We study how well deep neural networks approximate and learn spectral Barron functions. Recent studies have shown that these function classes can be efficiently approximated by shallow neural networks without suffering from the curse of dimensionality. We complement these results by providing new approximation bounds f...

Songqiu Ma, Yunfei Yang · 0 citations
#machine learning Preprint Sep 2026

Warm-starting PDE solvers with any-dimensional machine learning

Any-dimensional machine learning models, such as graph neural networks (GNNs), can be naturally trained and evaluated on inputs of different sizes and dimensions. Inspired by the GNN transferability literature, we show mathematical conditions under which a partial differential equation (PDE) learning-based solver can b...

Wilson G. Gregory, G. Kevrekidis, Ben Blum-Smith et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Sharp Statistical Rates for Asynchronous TD Learning with Markovian Data

We study the last iterate of standard tabular temporal-difference (TD) learning from a single trajectory of a finite Markov reward process. For discount factor $\gamma$, write $H=(1-\gamma)^{-1}$, and let $\mu_{\min}$ and $t_{\operatorname{mix}}$ denote the minimum stationary probability and total-variation mixing time...

Yang Peng · 0 citations
#machine learning Preprint Sep 2026

Optimal VC Dimension of Contrastive Learning with Margin

Contrastive learning is a successful paradigm for learning $d$-dimensional geometric representations from a collection of ``anchor--positive--negative''triplets $(i,j^{+},k^{-})$, indicating that ``item $i$ is closer to $j$ than to $k$.''Despite its success, understanding why contrastive learning leads to representatio...

Dionysis Arvanitakis, Vaggos Chatziafratis, Yi-Pei Luo et al. · 0 citations
#machine learning Preprint Sep 2026

Generative sequence modeling for infinite memory processes via predictive states

We consider estimating the one-step-ahead conditional distribution of a multivariate stochastic process. Many existing approaches rely on assumptions such as finite-range memory, sparsity, or additivity, which can be poorly suited to processes with long-range nonlinear interactions. However, without such structural ass...

Michael Wieck-Sosa, C. Shalizi · 0 citations
#machine learning Preprint Sep 2026

Advantage of Sample Complexity in Quantum PAC Learning Requires Inverse Access to State-Preparation Unitaries

Whether quantum computation can reduce the amount of data sampled from an unknown probability distribution required to learn a prediction rule is a fundamental question in quantum machine learning. Quantum PAC learning studies this question using quantum data as a quantum state whose squared amplitudes encode the unkno...

Natsuto Isogai, Satoshi Yoshida, M. Murao · 0 citations
#machine learning Preprint Sep 2026

Lower Bounds for Linear-Oracle Online Learning

Can a constant number of linear minimizations per round improve on the $T^{3/4}$ regret rate of online Frank-Wolfe on general convex sets? Weibel et al. conjectured that fixed-coefficient methods cannot. We prove their conjecture and extend the lower bound to every deterministic learner in an oracle-only model. The lea...

Mohit Sinha · 0 citations
#machine learning Preprint Sep 2026

Robust LassoNet: Enhancing Feature Selection in Neural Networks via Robust Loss Functions

Feature selection in neural networks remains a challenging problem, particularly in the presence of noisy or contaminated data. LassoNet is a recent approach that addresses this issue by combining neural networks with hierarchical sparsity constraints, enabling simultaneous prediction and variable selection. However, i...

D. De Canditiis, I. De Feis, P. Stolfi · 0 citations
#machine learning Preprint Open access Oct 2026

Conformal Factuality Control for Multi-Hop Retrieval-Augmented Generation

Retrieval-augmented generation (RAG) can ground large language models in external evidence, but retrieved context does not guarantee that generated claims are factually supported. This problem is especially relevant in multi-hop RAG, where retrieval and reasoning proceed through multiple dependent stages. We study whet...

Muhammad Aimal Rehman, Chi-Kuang Yeh · 0 citations
#machine learning Preprint Sep 2026

Distribution Matching Distillation for Continuous Diffusion Language Models

Continuous diffusion language models generate all tokens in parallel, yet high-quality generation can still require hundreds of network evaluations (NFEs). We study how distributional distillation can reduce this cost by exploiting the student's probabilistic token outputs. Our unified formulation connects the student'...

Paul Le Van Kiem, Dario Shariatian, Umut Simsekli et al. · 0 citations
#machine learning Preprint Sep 2026

Cheap to Draw, Expensive to Trust: Certifying Test-Time Scaling Curves

Sampling several answers and keeping the one a verifier scores highest is one of the simplest ways to buy accuracy at test time. Its effect is reported as a scaling curve: accuracy against the number $k$ of sampled answers. The curve is cheap to draw and expensive to trust. A budget read off it is chosen after looking...

S. Sarkar, S. Baichoo · 0 citations
#machine learning Preprint Sep 2026

Gromov-Wasserstein Distillation for Inductive Multi-View Embedding

Gromov-Wasserstein multidimensional scaling (GW-MDS) learns low-dimensional representations from relational data but remains transductive, providing no explicit mapping for unseen samples. We introduce an inductive framework based on barycentric distillation. A GW-MDS teacher learns a latent support and an optimal tran...

Rafael Pereira Eufrazio, Eduardo Fernandes Montesuma, C. C. Cavalcante · 0 citations

From tech blogs

See all →
Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.