Skip to content

Author

Hyeok‐Min Gwon

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access Jul 2026

Evaluating reinforcement learning from human feedback for task‐oriented dialogue systems

Reinforcement learning from human feedback (RLHF) has shown strong potential for aligning language models, but its role in task‐oriented dialogue (TOD) remains unclear. In TOD, models are typically trained with local turn‐level supervision, while system behavior is evaluated through broader interaction‐level properties. This mismatch becomes more challenging in online settings, where explicit dialogue‐level rewards and human preference annotations are unavailable. In this work, we study whether RLHF can be usefully applied to TOD under this limitation. We consider two task‐annotation regimes, partially annotated and fully annotated TOD data, and construct pseudo‐preference pairs using empirical ranking heuristics motivated by prior work on synthetic feedback and model‐based ranking signals. We then train reward models on the constructed pairs and optimize dialogue policies with Preference policy optimization (PPO) using simulator‐generated online trajectories. Experiments on MultiWOZ 2.1 show that the proposed RLHF approach consistently improves corpus‐based evaluation over supervised baselines, while simulator‐based effects remain mixed and backbone‐dependent.

Hyeok‐Min Gwon, Yohan Lee, Jin-Xia Huang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.