Skip to content

MaD-RL: Matching Distributions for Calibrating LLMs with Reinforcement Learning

Sep 2026 · 0 citations · 33 references
Computer Science

TL;DR

This work proposes a general RL-based framework for Distribution Matching allowing matching the distribution of a latent categorical attribute of model outputs to a specified target distribution and proposes reward functions for other divergences such as KL and Jensen-Shannon and motivate them with theoretical justification.

Abstract

Reinforcement learning (RL) is widely used in language-model post-training to maximize rewards assigned to individual model outputs, such as scores from binary verifiers or reward models trained on human feedback. However, applications such as synthetic-data generation, fairness-related constraint satisfaction, and policy exploration require controlling the distribution of outputs across model generations rather than only maximizing expected reward. We propose a general RL-based framework for \textit{Distribution Matching} allowing matching the distribution of a latent categorical attribute of model outputs to a specified target distribution. Empirically, we demonstrate that dominant post-training recipes such as Group Relative Policy Optimization (GRPO) reduce output diversity by concentrating policy probability towards a single mode. Entropy regularization and sampling temperature can improve the spread of the distribution but have constrained effectiveness, limited to apply only in token space and toward uniform distributions. We show that prior work in this area is a specific case of Distribution Matching involving the $L_2$ divergence. We then propose reward functions for other divergences such as KL and Jensen-Shannon and motivate them with theoretical justification. Finally, we demonstrate the effectiveness of our approach on a set of experiments involving mathematical reasoning and programming.

View source

Similar papers

Preprint Aug 2026

Best Practice Critic Optimization

Best Practice Critic Optimization (BPCO) is developed, a recipe that combines DPPO, value predictions bounded to the reward range, Monte Carlo value targets, unnormalized policy advantages, and length-adaptive generalized advantage estimation and shows that a carefully designed critic provides a reliable alternative to...

Penghui Qi, Xiang-Xin Zhou, Wee-Sun Lee · 2 citations · ⚡1
#artificial intelligence Preprint Sep 2026

Rethinking Training-Inference Mismatch in LLM Reinforcement Learning: Where It Arises and How to Correct It

We study training-inference mismatch in reinforcement learning with verifiable rewards (RLVR) for large language models, where rollouts are sampled by an inference engine while gradients are computed by a training engine, and the two engines assign different probabilities to the same tokens. To account for this discrep...

Tian-Run Yu, Kai-Xiang Zhao, Shang-Zhe Li et al. · 0 citations
#artificial intelligence Preprint Sep 2026

PACT: From Credit Assignment to Critic Alignment

Policy Aligned Critic Training (PACT), which adopts an Actor-then-Critic update order to apply importance sampling correction to critic training and better align the critic with the updated policy, is developed.

Jia-Yan Fu, Hang Xu, Yong Zhang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Timestep Weighting: A Hidden Key to Effective ELBO-Based Flow-Matching RL

ELBO-based reinforcement learning offers a sampler-agnostic approach to fine-tuning flow matching models with reward feedback. Timestep weighting in ELBO-based RL has large impact on performance, and it also provides a unified view (as we show in this work) to understand prediction losses heuristically chosen in prior...

Qin-Wei Ma, Jing-Zhe Shi, Si-Min Fan et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.