Preprint
Aug 2026
SecOPD: Mitigating Adaptive Prompt Injections by On-Policy Distillation
This paper proposes Secure On-Policy Distillation (SecOPD) that provides token-level feedback to guide defensive fine-tuning, and generalizes to domains completely unseen in training.
Yibo Peng, Long Lian, David A. Wagner et al.
· 0 citations