With recent advancements in large language models (LLMs) and LLM-based agents, these agents are becoming increasingly autonomous and gaining broader access to act on users'behalf on the internet. However, the vulnerability of automated agents deployed on social media platforms (e.g., for managing a user's personal acco...
Real-time payment fraud detection is a non-stationary streaming prediction problem: adversaries adapt before supervised labels mature, and localized burst attacks can cause losses before retraining. Production systems typically rely on tabular classifiers and rules, which can struggle to capture these emerging sequenti...
Many Security Operations Centers rely on signature-based Network Intrusion Detection Systems like Suricata, yet detection rule engineering remains understudied. We investigate this process by introducing SuriCap, a platform for rule engineering exercises, and hosting CTF-style workshops where 60 participants, trained M...
Koen T. W. Teuwen, Emmanuele Zambon, Luca Allodi· 0 citations
The tremendous commercial potential of large language models (LLMs) has heightened concerns over their unauthorized use. To address this, we focus on the task of identifying the origin of black-box LLMs. We further propose PlugAE, an effective and efficient identification method that proactively leverages LLM-specific...
Ziqing Yang, Yixin Wu, Yun Shen et al.· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
This work develops a method for optimizing the auditor's canary set to improve privacy auditing, leveraging recent work on metagradient optimization and demonstrates that in certain instances, using such optimized canaries can improve empirical lower bounds for differentially private image classification models by seve...
Matteo Boglioni, Terrance Liu, Andrew Ilyas et al.· arXiv.org· 7 citations· ⚡1
Concept drift, driven by the rapid evolution of Android malware, severely degrades the performance of machine learning detectors. Current adaptation strategies are often reactive, responding only after performance has dropped and imposing a significant manual annotation burden, or they are proactive but rely on unstabl...
Han Chen, Han-Chen Wang, Hong-Mei Chen et al.· 0 citations
Evaluation of large language models (LLMs) for safety, security, and privacy (SSP) relies heavily on static benchmarks, which suffer from score saturation, data contamination, and aggregation artifacts, and fail to capture sensitivity to linguistic variation. As a result, models that perform well on fixed test sets oft...
Fatih Deniz, Yazan Boshmaf, Issa Khalil· 0 citations
Rigorous expert validation confirms seed plausibility and realism for meaningful LLM safety evaluation, and provides an expert-validated, finance-specific rubric that goes beyond disclaimer checks, aligns more closely with human experts than static one-size-fits-all rubrics, and reduces critical false negatives from 28...
Chae-mi Kim, Daeyoung Park, Junghwan Kim et al.· arXiv.org· 1 citation
Verifier reliability is not portable across tasks. A verifier-guided self-DPO pipeline with genuine held-out gains on MathVista (+9.6 points on self-training data, +8.0 held out) can be harmful on MMMU. The failure is invisible from the target-task self-training signal: over six learner-verifier configurations, MMMU se...
Adaptive Unlearning is presented, a post-deployment framework that surgically suppresses package hallucinations while preserving general model utility and relies entirely on model-generated data and requires no human annotation, representing a post-deployment hallucination mitigation framework.
Joseph Spracklen, Pedram Aghazadeh, F. Koushanfar et al.· 0 citations
This analysis indicates that robust Long-Term Memory security cannot be retrofitted at retrieval or execution time alone, but must be anchored in storage-time provenance, versioning, and policy-aware retention from the outset.
Zehao Lin, Xixuan Hao, Renyu Fu et al.· 20 citations
This paper presents the first systematic audit between official LLM APIs and corresponding shadow APIs, and uncovers widespread behavioral inconsistency and fingerprint-based evidence consistent with deceptive model claims in a subset of audited endpoints.