GoodLiars: A Multi-Turn Extension of Reinforcement Learning-Based Belief Disruption
This work extends GOODLIAR from a single-turn attack to a multi-turn one and test two algorithms: a multi-turn variant of GOODLIAR’s DA-ILQL, an offline RL method with on-policy data aggregation, and a Proximal Policy Optimization (PPO) attacker, which is adopted because its per-turn rewards better match a multi-turn conversation.