Hypergraphs record multi-way interactions among entities. Extracting information from the combinatorial structure underlying observed multi-way interactions is a central task in many real-world problems. Existing methods face several limitations. First, many deep architectures for hypergraphs do not explicitly exploit...
Zi-Meng Li, Shi-Hao Wu, Gong-Jun Xu et al.· 0 citations
We introduce a sampling approach for energy- and score-based generative models that requires no gradient evaluations of the model. Replacing the drift term that would normally contain the score $\nabla_\mathbf{x} \log p_\theta(\bf{x})$ with a high-frequency dithered cosine of the model's \textit{value}, $\sqrt{\alpha\o...
Simulation-free training of latent Stochastic Differential Equations (SDEs) relies on a variational posterior process whose one-time marginals are tractable, typically Gaussian. Such marginals, however, do not determine the underlying dynamics: many processes share the same marginals while differing in their temporal s...
Grigory Bartosh, Christian A. Naesseth· 0 citations
Sparse causal discovery calls for methods that exploit graph structure without estimating high-dimensional densities. We introduce Fisher Information Completion Search (FiCS), a source-first algorithm for additive noise models that uses one local Fisher score for both ordering and parent selection. Under regularity and...
Byeongguk Kang, Donghyeon Lee, Euijong Song et al.· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
Graph Neural Networks (GNNs) often struggle to capture long-range dependencies due to over-squashing -- a phenomenon in which the repeated compression of node embeddings into finite-size messages causes representations to collapse. Over-squashing is most often diagnosed as a property of the graph topology, with effecti...
Andr\'e Ribeiro, Germano Barcelos, Amauri H. Souza et al.· 0 citations
Applications such as engineering design often require us to optimize a black-box function, i.e., a system whose inner processing is not analytically known and whose gradients are not available. Practitioners often have a fixed budget for the number of function evaluations and the performance of an optimization algorith...
We propose a data-driven framework for estimating the Bayesian Cram\'er-Rao bound (CRB) in high-dimensional imaging systems with complex, analytically intractable priors. Direct CRB computation is challenging in this setting due to the need to model the prior score and to form and invert the Bayesian Fisher information...
Evan Scope Crafts, Thomas Wynn, Seonyeong Park et al.· 0 citations
Probability-flow ordinary differential equations (PF-ODEs) are widely used as deterministic samplers for score-based diffusion models. Their usual justification is that the Fokker--Planck equation of a diffusion can be rewritten as a continuity equation driven by the score function of the forward diffusion. This identi...
Treatment efficacy is traditionally demonstrated on the basis of a single primary outcome. However, clinical decision-making usually requires consideration of multiple outcomes, balancing expected benefits against potential risks. The relative value assigned to these outcomes varies substantially from one patient to an...
Lola Giordani, Mathieu Even, Chlo\'e Geoffroy et al.· 0 citations
Neural network-based approaches have emerged as efficient alternatives to traditional optimization-based procedures for the calibration of stochastic volatility models. However, existing work has focused primarily on predictive accuracy, with comparatively little attention devoted to understanding the structure of the...
Sha\"in Afzali, Serena Della Corte, Antonis Papapantoleon· 0 citations
We study online control of a known linear dynamical system with adversarial costs and bounded disturbances, measuring regret against a general class of benchmark policies. We introduce counterfactual tracking, which separates the challenge of learning from the challenge of controlling the system. An online learner buil...
Modern artificial intelligence (AI) systems are powered by foundation models. This paper presents a new set of foundation models, called Llama 3. It is a herd of language models that natively support multilinguality, coding, reasoning, and tool usage. Our largest model is a dense Transformer with 405B parameters and a...
Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity. The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.