Robust Context-Aware Detection of Malicious Instructions in Text
The proposed approach for malicious sentence classification that is both context- and query-aware and outperforms state-of-the-art IPI defense baselines under static attacks, while in the case of adaptive attacks, the AT variants provide significantly higher utility, lower attack success rate, and often both.