Skip to content

Category

data science

2,430 papers

#machine learning Preprint Oct 2026

Iterating Consistency Models: Stability, Error Bounds and Noise Schedules

Consistency models (CMs) have become a leading approach for generating high-quality samples in few steps. However, adding steps can improve or degrade sample quality in ways that are highly sensitive to the schedule and that existing theory does not fully explain. To provide accuracy guarantees and guide CM sampler des...

Alessio Spagnoletti, A. Haji-Ali, Andrés Almansa et al. · 0 citations
#machine learning Preprint Oct 2026

DAWIS: Data Assimilation with Windowed Inverse Sampling via Multitask Interpolants

Flow- and diffusion-based generative models have recently emerged as flexible and highly efficient forecasting models for dynamical systems. When combined with inference-time guidance, they offer a promising route to high-dimensional non-Gaussian data assimilation (DA), the problem of combining forecasts with observati...

Erik Wikingsson, Martin Andrae, Tomas Landelius et al. · 0 citations
#machine learning Preprint Open access Oct 2026

SDECast: Probabilistic Weather Forecasting in Continuous Time with Neural SDEs

Existing machine learning weather forecasting models typically generate forecasts through autoregressive rollouts at a fixed temporal resolution. While highly efficient for long-range prediction, this formulation can suffer from severe error accumulation when used with shorter time steps and does not explicitly encode...

Maria Marchenko, Martin Andrae, Fredrik Lindsten et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Near-Optimal Convex Optimization with Lazy Second-Order Oracles

This paper studies the complexity of convex optimization using lazy second-order oracles (Doikov, Chayti, and Jaggi, ICML 2023), where an algorithm queries gradients every iteration and Hessians once per $m$ iterations. Under this setting, we show a lower bound of $\Omega(m+ m^{1/7} \epsilon^{-2/7})$ on the number of t...

Xinliang Zhang, Lesi Chen, Chengchang Liu et al. · 0 citations
#machine learning Preprint Oct 2026

Predictively Oriented Gaussian Process Posteriors

Gaussian Processes (GPs) are a powerful tool for modelling and quantifying uncertainty in functional relationships. However, they require practitioners to make a number of design decisions, such as the choice of the kernel and the observation model. Suboptimal choices can produce misspecified models that do not capture...

Callum Lau, Jeremias Knoblauch, Louis Sharrock · 0 citations
#machine learning Preprint Open access Oct 2026

Invariance of Clustering Operations in Causal Effect Identification

Clustering variables in causal graphs reduces the size of the graph and simplifies causal inference. However, arbitrary clustering can alter crucial causal relations among variables and lead to erroneous conclusions. While the identifiability of a causal effect in the clustered graph implies the identifiability in the...

Jani Nyk\"anen, Otto Tabell, Santtu Tikka et al. · 0 citations
#machine learning Preprint Oct 2026

GTDD: Generative Test-Driven Development for AI Coding Agents with Adversarial Testing

Test-driven development gives AI coding agents executable requirements for implementing software. Because these agents can adapt their implementations to the examples they observe, passing a predetermined collection of tests can leave substantial parts of the intended behavior unimplemented. We propose Generative Test-...

Masahiro Kato · 0 citations
#machine learning Preprint Open access Oct 2026

Hold-Out Scoring for Efficient Gaussian DAG Learning

High-dimensional Gaussian DAG learning faces a statistical-computational gap: methods with sharp sample complexity rely on computationally expensive subset search and a supplied indegree bound, whereas polynomial-time alternatives have less favorable sample complexity. We introduce HOST, an efficient DAG learning algor...

Donguk Shin, Byeongguk Kang, Inseol Lee et al. · 0 citations
#machine learning Preprint Oct 2026

High-Dimensional Asymptotics and Dataset Selection for Private Transfer Learning

To commit to buying external data or participate in collaborative learning, one must decide whether the additional data will improve prediction enough to justify the cost. This comes with several challenges: (i) the decision often relies only on aggregated statistics available publicly, rather than individual-level dat...

Filip Kovacevic, Edwige Cyffers, Stefano Sarao Mannelli et al. · 0 citations
#machine learning Preprint Open access Oct 2026

ENCORE: Exact Non-equilibrium COntrol with Replica Exchange for Diffusion Generation

Inference-time control steers a pretrained generative model towards a target distribution without retraining. We study tilted targets $\pi_0\propto G_0\,p_0$, where $p_0$ is the sampler output distribution and $G_0$ is an evaluable reweighting function. Existing approaches rely on sequential annealing with sequential M...

Jiahao Yu, Saifuddin Syed, Jos\'e Miguel Hern\'andez-Lobato et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Conformal Prediction for Time Series with Deep Sequence Models

Recent advances in deep learning for time series prediction have amplified the need for reliable uncertainty quantification. Conformal prediction has gained attention as a distribution-free framework for constructing prediction intervals with coverage guarantees. However, its coverage guarantees rely on data exchangeab...

Junghwan Lee, Jonghyeok Lee, Yao Xie · 0 citations
#machine learning Preprint Open access Oct 2026

Expected Utility Regret Rule: Minimax and Bayes Optimal Portfolio Choice

This study considers the problem of portfolio choice, where we recommend a portfolio to an investor to maximize the expected utility of their wealth. Our goal is to construct an asymptotically optimal portfolio choice rule in terms of expected utility regret, the difference between the expected utility of an oracle inv...

Masahiro Kato · 0 citations

From tech blogs

See all →
Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.