AEGIS: Real-Time Latent-Space Backdoor Detection for Dependable and Secure Small Language Model Inference
Backdoor attacks pose a serious threat to small language models (SLMs) because compromised models can behave normally on benign inputs while producing attacker-specified outputs when a hidden trigger is activated. Existing defenses commonly require model retraining, operate only before deployment, rely on input-level s...