Detecting and Localizing Segment-Level Poisoning in Multi-Source LLM-Agent Inputs
ActProbe, an internal-state-based framework for detecting and localizing poisoned segments in multi-source LLM inputs, is proposed and remains effective against defense-aware adaptive attacks and can protect black-box APIs through surrogate-based poisoned-segment removal.
Xue Tan, Chang-Hui Wang, Sanrui Yang et al.
· 0 citations