Skip to content

ORPG: Reconciling Multiple Reward Objectives through Objective-wise Policy Gradients

Sep 2026 · 0 citations · 35 references
Computer Science

TL;DR

Objective-wise Reconciled Policy Gradient (ORPG), which constructs a separate clipped policy objective for each reward and reconciles the resulting gradients into one policy update, achieves the highest average full-budget accuracy and three-budget hypervolume among the compared methods.

Abstract

Multi-reward policy optimization requires a joint update that reflects both the learning signals and the intended relationships among objectives. We introduce Objective-wise Reconciled Policy Gradient (ORPG), which constructs a separate clipped policy objective for each reward and reconciles the resulting gradients into one policy update. For compatible gradients, a cosine-dependent interpolation coordinates their contributions through a partially normalized reference while preserving the norm of their sum. We characterize this update as the unique solution of a spherical directional compromise. For conflicting gradients, projection follows the task's priorities. We evaluate the same compatible rule in helpfulness--safety alignment and correctness--cost optimization for mathematical reasoning. ORPG substantially improves average Useful and Harmless scores over the strongest external baseline on each axis. In mathematics, it achieves the highest average full-budget accuracy and three-budget hypervolume among the compared methods, with more accurate and shorter responses than the initial policy. Component comparisons and training dynamics show the larger contribution of compatible coordination and a complementary benefit from conflict handling. These results support gradient reconciliation for objectives with equal standing and for objectives with an explicit priority.

View source

Similar papers

Preprint Aug 2026

Learn What's Left, Not What's Mastered: Saturation Aware Advantage Reweighting for Multi-Reward Policy Optimization

Saturation Aware Advantage Reweighting for Multi-Reward Policy Optimization (SA-MRPO), which standardizes each reward objective independently and adaptively discounts its contribution according to a batch-level estimate of objective saturation, and dynamically reallocates optimization effort toward under-optimized obje...

Yi-Xuan Wang, Yi-Fei Chen, Haichao Zhang et al. · 2 citations · ⚡1
#machine learning Preprint Sep 2026

IncentRL: The Trade-Off Between Preference Guidance and Task Performance

Preference-based reward shaping can guide reinforcement learning, but adding preference signals to the reward may unintentionally change the task being optimized. We address this problem with IncentRL, a framework that introduces preference guidance while explicitly characterizing its effect on external-task performanc...

Xue-Ning Wu, Yan-Lan Kang, Shen Yin · 0 citations
#machine learning Preprint Sep 2026

When Sparse Reward Meets Dense Distillation: Training Dynamics of On-Policy Distillation

The cross-signal NTK is introduced, a token-level statistic that measures the alignment between reward and distillation gradients at position n and an empirical threshold beyond which naive mixing can lead to persistent training collapse is revealed.

Xin-Ke Jiang, Tao Feng, Zhi-Bang Yang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

A Better Spur Should Start From Each Objective

This work proposes Multi-Marginal Preference Optimization (MMPO), a fine-grained framework that intervenes at the data, gradient, and constraint levels rather than relying on coarse-grained global scalarization to address optimization conflicts among multiple objectives in real-world deployment scenarios.

Shang-Wen Mao, Hao Zhang, Guangtao Nie et al. · 1 citation · ⚡1
Preprint Aug 2026

I-SDPO: Instance-Level Adaptive Self-Distillation Policy Optimization

I-SDPO (Instance-Level Adaptive Self-Distillation Policy Optimization), which treats teacher reliance as capability-dependent and uses imitation only where group-relative rewards are uninformative, obtains the best result in all four scientific domains.

Yubo Zhang, Xin-Hong Ma, Zezhong Tan et al. · 1 citation

Related blog posts

MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.