Skip to content
Open access

Rapid Bayesian computation and estimation for neural networks via log-concave coupling

Nov 2024 · Mathematical Statistics and Learning · 1 citation · 66 references
Mathematics

Abstract

This paper presents the study of a Bayesian estimation procedure for single-hidden-layer neural networks using \ell_{1} controlled neuron weight vectors. We study the structure of the posterior density and provide a representation that makes it amenable to rapid sampling via Markov Chain Monte Carlo (MCMC). Let the neural network have K neurons with internal weights of dimension d and fix the outer weights. Thus, there are Kd parameters overall. With  N data observations, use a gain parameter or inverse temperature of \beta in the posterior density for the internal weights.The posterior is intrinsically multi-modal and not naturally suited to rapid mixing of direct MCMC algorithms. For a continuous uniform prior on the \ell_{1} ball, we demonstrate that the posterior density can be written as a mixture density with suitably defined auxiliary random variables, where the mixture components are log-concave. Furthermore, when the total number of model parameters Kd is large enough that Kd \geq C(\beta N)^{2} , the mixing distribution of the auxiliary random variables is also log-concave. Thus, neuron parameters can be sampled from the posterior by only sampling log-concave densities. The authors refer to the pairing of weights with such auxiliary random variables as a log-concave coupling.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.