This work interprets chain-of-thought reasoning as a latent variable modeling problem and demonstrates that this distribution-matching paradigm of LLM fine-tuning can serve as an effective alternative to maximum-likelihood training and reward-maximizing policy optimization.
Edward J. Hu, Moksh Jain, Eric Elmoznino et al.· International Conference on...· 110 citations· ⚡19
This paper builds bridges between two families of probabilistic algorithms: (hierarchical) variational inference (VI), which is typically used to model distributions over continuous spaces, and generative flow networks (GFlowNets), which have been used for distributions over discrete structures such as graphs. We demonstrate that, in certain cases, VI algorithms are equivalent to special cases of GFlowNets in the sense of equality of expected gradients of their learning objectives. We then point out the differences between the two families and show how these differences emerge experimentally. Notably, GFlowNets, which borrow ideas from reinforcement learning, are more amenable than VI to off-policy training without the cost of high gradient variance induced by importance sampling. We argue that this property of GFlowNets can provide advantages for capturing diversity in multimodal target distributions.
Esmeralda S. Whitammer, S. Lahlou, T. Deleu et al.· International Conference on...· 120 citations· ⚡9
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.