The teacher, an auxiliary behavior model, is trained to sample high-loss regions of the student and can generalize across unexplored modes, thereby enhancing mode coverage by providing an efficient training curriculum.
Minsu Kim, Sanghyeok Choi, Taeyoung Yun et al.· International Conference on...· 27 citations· ⚡6
Amortized sampling of the posterior over data is studied, and the asymptotic correctness of a data-free learning objective, relative trajectory balance, is proved for training a diffusion model that samples from this posterior, a problem that existing methods solve only approximately or in restricted cases.
S. Venkatraman, Moksh Jain, Luca Scimeca et al.· Neural Information Processin...· 75 citations· ⚡5
It is proved that any global minimizer of the trajectory balance objective can define a policy that samples exactly from the target distribution, and empirically demonstrate the benefits of the trajectories balance objective for GFlowNet convergence, diversity of generated samples, and robustness to long action sequences and large action spaces.
Esmeralda S. Whitammer, Moksh Jain, Emmanuel Bengio et al.· Neural Information Processin...· 302 citations· ⚡60
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.