Causal discovery in time series is increasingly performed using nonlinear machine-learning models, yet the resulting causal relationships are almost always summarized by scalar edge scores. We argue that this practice obscures the true object learned by nonlinear autoregressive models: a state-dependent function whose...
Valentina V. Kuskova, Dmitry Zaytsev, Michael Coppedge· 0 citations
Physics-informed machine learning (PIML) integrates mechanistic knowledge, typically through partial differential equations (PDEs), into data-driven models. Despite strong empirical performance, its statistical generalisation properties remain poorly understood, especially for regression with unbounded losses. We devel...
Thien V. Nguyen, Amaury Habrard, Benjamin Guedj· 0 citations
Spectral partial least squares (PLS-SVD) estimates the directions shared by two views of the same samples. We analyze it using a rank-one regression model, with entries of both views missing completely at random and filled with zeros. Missing response entries weaken the signal, whereas missing design entries also tilt...
Anders Gj{\o}lbye, Emma Kargaard, Ida Kargaard et al.· 0 citations
Deep equilibrium models (DEQs) have recently emerged as a powerful paradigm for training infinitely deep weight-tied neural networks that achieve state of the art performance across many modern machine learning tasks. Despite their practical success, theoretically understanding the gradient descent dynamics for trainin...
Sanjit Dandapanthula, Aaditya Ramdas· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
Inverse problems governed by partial differential equations (PDEs) arise widely in science and engineering, but are often ill-posed and limited by sparse, noisy observations. In many applications, measurements reveal only part of the physical state, while the domain geometry may also be unknown even though it shapes th...
Yajie Ji, Sifan Wang, Zhikai Wu et al.· 0 citations
We propose a novel distributionally robust $Q$-learning algorithm for the non-tabular case accounting for continuous state spaces where the state transition of the underlying Markov decision process is subject to model uncertainty. The uncertainty is taken into account by considering the worst-case transition from a ba...
Chung I Lu, Julian Sester, Aijia Zhang· 0 citations
I show that ordinary least squares (OLS) predictions can be rewritten as the output of a restricted attention module, akin to those forming the backbone of large language models. The connection comes from viewing OLS as a similarity-based prediction rule in a learned embedding space. In this representation, least squar...
Some diffusion posterior samplers construct Gaussian-tilted intermediate distributions along the reverse process. We observe that these targets can be pulled back to clean-space posteriors with weaker conditioning, with samples transported analytically to the corresponding noisy-space target through a Gaussian bridge....
A recurring design principle in modern optimizers is to decouple update magnitude from the raw gradient norm, yet its consequences for learning-curve and resource scaling remain unclear. We isolate this mechanism by studying normalized SGD in a random-feature model with power-law teacher and data covariance. Fixed-norm...
CoVariance Neural Networks and their extensions have emerged as effective tools for processing multivariate data, deriving graph shift operators directly from second-order statistics. These architectures, however, are designed for independent and identically distributed observations and do not fully capture the joint s...
Lorenzo Marinucci, Leonardo Di Nino, Gabriele D'Acunto et al.· 0 citations
We study contextual bandits in which a surrogate is observed after the action but before the learner decides whether to acquire the primary outcome that defines action value and regret. The value of acquiring the primary outcome depends on both decision relevance (how much the current outcome matters for comparing poli...
Evaluating text generation requires measuring how well the generated distribution matches the data distribution. For autoregressive models, this is done by the perplexity. Diffusion and flow-based language models can only provide a likelihood bound, whose tightness differs between model families. Sample-based substitut...
Gregor Kornhardt, Moritz Piening, Jannis Chemseddine et al.· 0 citations
Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity. The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.