Before a machine learning model ships, it often has to pass a suite of automated tests. Requiring every test to pass looks safe, yet it can reject many models that would have served users well, and it does not say how trustworthy a passing model actually is. We treat the release gate as a design problem: choose how man...
We study an endogenous nonstationary stochastic bandit problem with latent linear dynamics, where actions affect both immediate rewards and the future evolution of an unobserved latent state. Rewards are bilinear in the current action and latent state, inducing history-dependent rewards and a nontrivial long-horizon pl...
Taehyun Hwang, Hyunjun Choi, Heesang Ann et al.· 0 citations
We consider the estimation of drift, diffusion, and noise covariance from discrete observations of stochastic differential equations driven by Gaussian processes. For a fixed observation horizon $T>0$ and a known initial state $x_0\in\mathbb R$, we study \begin{equation*}
dX_t=a(X_t)\,dt+\sigma(X_t)\,dZ_t^{\beta,f},...
There is considerable interest in using time-varying electricity prices to shape consumer demand response, and better align energy demand with renewable production. However, optimal prices generally vary over time in response to complex signals such as weather forecasts, sunrise/sunset times, and day-of-week patterns;...
Jing Shang, Mohammad Mehrabi, Xinyang Zhou et al.· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
Uncertainty quantification is critical in scientific machine learning, where black-box, image-based models are increasingly deployed in high-stakes settings. In many such applications, model outputs inform costly decisions, yet most methods provide only point estimates without quantifying predictive uncertainty. This c...
Carrie J. Lei-Cramer, Michael S. Jones, Laura J. Wendelberger· 0 citations
Neural operators learn maps between function spaces, while hereditary network dynamics are described by Volterra resolvents with non-rational Laplace symbols. We introduce a fractional Laplace neural operator (fLNO) that embeds this structure in the learned map. For commuting excitation--Laplacian pairs, one block grap...
When minimizing the squared-error loss, the popular CP decomposition can be interpreted as parameter inference in a Gaussian model with a low-rank mean tensor and constant variance across the tensor entries. We introduce heteroskedastic-CP (HCP), which models entrywise variability with a non-constant, low-rank precisio...
Continual parameter-efficient fine-tuning for large language models (LLMs) must balance retention of previously acquired knowledge, adaptation to new tasks, and strict parameter budgets. We present \textbf{ChainLoRA}, a replay-free continual merging framework built on chain-updated task-vector geometry. From a paramete...
Hang Yin, Hao-Zhe Wang, Yu-Hua Luo et al.· 0 citations
A weight space network (or metanetwork) takes the weights of another neural network as input and predicts properties of it. Most prior work trains such models on input networks of one or a few fixed sizes and evaluates them in-distribution. The few attempts at out-of-distribution size generalization remain limited in s...
Yuxin Ma, Adir Dayan, Yam Eitan et al.· 0 citations
How much risk does a small reweighted training support retain? For finite weighted least squares with the minimum-norm learner, we prove the exact law $\Gamma_d(n)=3-n/d$ throughout $\lceil3d/2\rceil\leq n\leq2d-1$. The guarantee covers every observed feature rank and uses selections that preserve the full feature span...
We consider the problem of sampling from Gibbs distributions on matrix spaces whose potential energies are neither convex nor globally gradient-Lipschitz. We introduce a family of non-quadratic kinetic energies that lead to a new underdamped Langevin system with momentum preconditioning, in which the gradient of the ki...
Decision-focused learning for linear optimization is complicated by the discontinuity of the optimizer, where small cost errors may leave the decision unchanged or move it to a different vertex. We show that this non-smooth pointwise behavior becomes locally quadratic after averaging over the data distribution, and we...
Konstantinos Ziliaskopoulos, A. Vinel, Alice E. Smith· 0 citations
Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity. The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.