Prompt Injection Detection for Email Agents Through Attack Chain Modeling
A detection framework that models this attack chain by combining a text detector, verifiers specific to each stage, explicit rule-based risk signals, user intent and action consistency analysis, and a logistic decision policy is proposed, which achieves a mean F1 score under the strict threshold setting policy.