#machine learning
Oct 2025
GREAT: Generalizable Backdoor Attacks in RLHF via Emotion-Aware Trigger Synthesis
This work develops GREAT, a novel framework for crafting natural distributional backdoors in RLHF, which targets harmful response generation for a vulnerable user subpopulation featured by semantically violent requests paired with emotionally angry triggers.
Subrat Kishore Dutta, Yuelin Xu, P. Pant et al.
· arXiv.org · 0 citations