ShalDeepDP: A Survey and Comparison of Data Poisoning in Statistical, Classical Deep Learning and Foundation Models
Advances in real-world adversarial threats have heightened concerns over the privacy and security of AI systems. One key threat is Data Poisoning, where malicious data is injected into training sets or model inputs to compromise model behavior. While much of the current research focuses on deep learning, particularly Foundation Models such as Large Language Models, there is limited understanding of how data poisoning affects different model eras. This survey presents a comparative analysis of data poisoning vulnerabilities in different model eras: statistical machine learning models, smaller-scale ”classical” deep learning models, and foundation models. Our goal is to inform practitioners of the trade-offs between robustness and performance when choosing models for different applications.