Skip to content

Category

data science

2,430 papers

#machine learning Preprint Open access Oct 2026

KESurv: A Kernel Ensemble Method for Patient-Specific Survival Prediction

Predicting patient-specific survival functions is crucial for clinicians in making informed decisions about patient care and treatment strategies. Among the various models available, the Survival Forest has demonstrated significant effectiveness in numerous scenarios. In this work, we propose an ensemble method that le...

Rahul Goswami · 0 citations
#machine learning Preprint Open access Oct 2026

Latent Similarity Gaussian Processes: A Theory-Grounded Approach to Personalized Suicide-Risk Forecasting for Clinical Decision-Support

Forecasting suicide risk is difficult due to the high heterogeneity of patients and the low base rate of suicide-related events (SREs). We present Latent Similarity Gaussian Processes (LSGPs), which embed patients in a continuous latent space to jointly model similarity and forecast risk. By selectively drawing informa...

Yaniv Yacoby, Weiwei Pan, Hope Neveux et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Watermarking: from Impossibility to Auditable Compliance

Article 50 (2) of the EU Artificial Intelligence Act requires providers of generative systems to make synthetic outputs machine-readable and detectable, while qualifying the effectiveness, interoperability, robustness, and reliability by technical feasibility, cost, content-specific limits, and the state of the art. Fo...

Fernando Delbianco, Fernando Tohm\'e, Hugo Acciarri · 0 citations
#machine learning Preprint Oct 2026

Finite-Sample Distribution Theory and Efficient Large-Scale Inference for Online Quantile Regression

This paper studies online quantile regression for large-scale and streaming data using Stochastic SubGradient Descent (SSGD) with constant learning rates. Classical offline inference for quantile regression is computationally and memory intensive. Existing works of online inference for quantile regression provide only...

Zi-Yang Wei, Jia-Qi Li, Lan Wang et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Isotropic Gaussian Processes Improve Vanilla Bayesian Optimization in High Dimensions

High-dimensional Bayesian optimization (BO) often fits Gaussian process (GP) surrogates from far fewer observations than input dimensions. Modern Vanilla BO can perform well in this regime with dimension-aware priors, initialization, and acquisition optimization, but it typically retains automatic relevance determinati...

Wei-Ting Tang, Madhav Muthyala, Joel A. Paulson · 0 citations
#machine learning Preprint Open access Oct 2026

Retrieval-Based In-Context Learning: A Domain Adaptation Framework

In-context retrieval (ICR) is a retrieval-based form of in-context learning (ICL) in which demonstrations are retrieved from a source database based on similarity to the query, rather than sampled independently. In this work, we formulate ICR as a type of domain adaptation problem, where the source distribution $P$ of...

Yilun Zhu, Naihao Deng, Yingcong Li et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Gaussian Limits for SGD Without Stationary Moments

Temporal dependence can separate the Gaussian approximation of stochastic gradient descent from its stationary moments. For unmodified least-squares SGD, we construct a design with standard Gaussian marginals whose stationary error has every positive moment infinite. Independent observations with the same marginals ins...

Xiaoli Li, Wei Biao Wu · 0 citations
#machine learning Preprint Oct 2026

Moment-Accurate Gaussian Mixtures for Constant-Step Stochastic Approximation

Local Gaussian models of constant-step learning predict output variability and expected losses, but weak convergence alone does not justify these moment predictions. We establish moment-accurate Gaussian mixtures by matching stationary energy with local Ornstein--Uhlenbeck limits, ruling out quadratic tail mass invisib...

Xiao-Lin Li, Wei-Biao Wu · 0 citations
#machine learning Preprint Oct 2026

Taylor Representations for Model-Free RL in Networked MDPs

In Networked Markov Decision Processes, transition dynamics are often unknown and the state--action space grows rapidly with the number of agents. In this setting, Taylor representations naturally approximate $Q$-functions, but a naive order-$n$ expansion over $N$ agents requires $\Theta(N^n)$ coefficients. We justify...

Salah Chikhi, Abdelhaq Chaoui, A. Ozdaglar et al. · 0 citations
#machine learning Preprint Open access Oct 2026

A Statistical Inference Framework for PMI Estimation and SGNS Word Embeddings

Pointwise Mutual Information (PMI) is a core measure of testing word association, and Skip-gram with Negative Sampling (SGNS) is essentially a method that implicitly factorizes a shifted PMI matrix. However, a systematic and well-rounded characterization of finite-sample uncertainty in PMI estimation remains absent and...

Zhongqi Fan · 0 citations
#machine learning Preprint Open access Oct 2026

Variance-Aware Fine-Grained Gap-Dependent Bounds for Online Reinforcement Learning

We study model-free online reinforcement learning (RL) for episodic tabular Markov decision processes, focusing on both gap-dependent regret and policy switching cost. While fine-grained gap-dependent analysis has been established for model-free RL algorithms using Hoeffding-type exploration bonuses, such results for m...

Haochen Zhang, Lingzhou Xue, Zhong Zheng · 0 citations
#machine learning Preprint Open access Oct 2026

Exact Fast Batch Simulation for Tabular Reinforcement Learning

Simulation is a fundamental computational primitive in reinforcement learning (RL), yet conventional simulation explicitly generates individual trajectories even when downstream procedures use only aggregate statistics. To address this, we develop an exact fast-simulation framework for finite-horizon tabular Markov dec...

Haochen Zhang, Lingzhou Xue, Zhong Zheng · 0 citations

From tech blogs

See all →
Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.