We study errors-in-variables problems in which a target study contains only a single error-prone surrogate of an unobserved exposure, while an external source study provides repeated surrogate measurements from a different population. The measurement error distribution is allowed to depend on the observed error-free va...
Multivariate functional data arise in many modern manufacturing systems, where multiple sensors record densely sampled process trajectories. Monitoring such data is challenging because nominal variation is strongly correlated across samples, sensors, and time, while faults may appear either as isolated deviations or as...
Mohammad N. Bisheh, Che-Yi Liao, Kamran Paynabar· 0 citations
This paper concerns how semantic context determines geometry in learned vector representations. Similarity is typically measured using cosine similarity, which provides a single fixed geometry. Semantic similarity, however, is inherently context dependent: two images may be similar because they depict the same object,...
Akash Kannan, Kiho Park, Victor Veitch· 0 citations
A central question in the prediction of stochastic processes is the extent to which past information can improve the probability of correctly predicting the next state. We introduce the Memory Prediction Excess (MPE) to address this question quantitatively. The MPE measures the average improvement in prediction accurac...
Jiahao Jiang· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
We introduce a geometric formulation of statistical feature learning for supervised regression. Feature learning is defined through a base--fiber decomposition: the base is the feature-side geometry produced by training, and the fiber is the learned feature space where estimation is performed. We prove this property fo...
Zong Shang, Tomoya Wakayama, Guillaume Lecu\'e et al.· 0 citations
We propose a method for choosing the shared memory parameter $\beta_1=\beta_2=\beta$ in Adam from a short pilot training. The selected $\beta$ remains fixed during the subsequent full training. A local model of Adam's normalized direction balances sampling variability against the delay introduced by averaging past grad...
Alberto Fern\'andez-Hern\'andez, Cristian P\'erez-Corral, Jose I. Mestre et al.· 0 citations
Reinforcement learning (RL) has become a standard tool for post-training language models on reasoning tasks, where the policy is updated by reward feedback while exploring the space of responses. Despite its empirical success, theoretical understanding of RL post-training remains limited, in particular of why on-policy...
We establish a polynomial sample complexity separation between symmetry-aware and symmetry-agnostic feature learning. We study growing-rank multi-index models with high-dimensional Gaussian covariates in $\mathbb{R}^d$ and $r=\Theta(d^\delta)$ teacher directions forming a cyclic symmetry orbit, where $0<\delta<1/2$. We...
Jivan Waber, Vanessa Piccolo, Yatin Dandi et al.· 0 citations
A mean shift between two data sources can be easy to detect but hard to remove without substantially changing their representations. We cast its removal as a statistical decision problem: from noisy differences between paired calibration measurements in $\mathbb{R}^d$, learn one linear map, applied to both sources unde...
Anuar Aimoldin, Yankai Chen, Ayana Mussabayeva et al.· 0 citations
Simulation-based inference (SBI) has become a powerful approach to Bayesian inference in complex scientific models whose likelihoods are difficult or impossible to evaluate. Amortized SBI learns reusable inference models from simulated data, enabling rapid posterior inference for new observations, and modern generative...
Xin Zhao, Nico Scherf, Robert Trampel et al.· 0 citations
Prior Fitted Networks (PFNs) such as TabPFN now rival established statistical procedures across prediction and estimation tasks. A natural explanation is that PFNs have the property of statistical adaptivity, that is, they perform nearly as well as a method tailored to the true data-generating model for a heterogeneous...
Martin Eppert, K. Balasubramanian, Subhro Ghosh et al.· 0 citations
Pathwise gradients are preferred for continuous random variables because they are unbiased, low variance, and work with a single sample. For discrete variables, however, the pathwise identity cannot generally be exact for every differentiable function. We propose a general framework to construct finite-order exact path...
Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity. The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.