Collaborative medical AI platforms allow researchers to train models on sensitive imaging data while restricting data export. However, trained models can serve as covert carriers of patient information: medical images may be encoded within model parameters and reconstructed outside the secure environment. Existing defe...
Elie Thellier, Hui-Yu Li, Nicholas Ayache et al.· 0 citations
Recent work suggests that obscure language registers can weaken large language model refusal behavior, especially when paired with black-box prompt optimization. It remains unclear whether failures stem from non-standard surface form, culturally grounded framing, or optimization over an expressive prompt-strategy space...
The Pareto frontier between the two defense objectives is characterized, and both rates are confirmed on seven binary-classification tasks spanning tabular, image, and language-model-feature inputs: label-plus-counterfactual access extracts the boundary with up to $200\times$ fewer queries than a published label-only b...
This work introduces artifact-mediated propagation, where adversarial content introduced through an artifact is stored in an assistant's persistent memory, reproduced in a subsequently created artifact, and acquired by another assistant that later reads it.
Sidharth Pulipaka, Anshu Sharma, Stanislau Hlebik et al.· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
Formation-consistent dispatch (FCD), which connects implementation analysis to execution authority, is presented, which produces provenance-bound over-approximations of declared in-scope effects from official source.
Geonwoo Kim, Brent ByungHoon Kang Korea Advanced Institute of Science, Technology· 0 citations
Veer is introduced, an agent-side runtime defense that leaves task planning to the base agent and intervenes on Web state when a proposed action would produce an unauthorized consequence, establishing task-relevant Web state as an effective runtime control target for protecting Web agents from deceptive outcomes.
DGF-Bench is a benchmark in which a board of agents (specialist gates and a General gate that consolidates their decisions) reviews synthetic dossiers while an attacker plants deceptive content in evidence the organization does not vouch for.
This work introduces PlanGuard, the first pre-execution detector that evaluates the physical safety of a complete multi-step plan in its current environment, and proposes Strong-Teacher Adaptive Compensation for On-Policy Distillation (STAC-OPD), which provides compact models with adaptive strong-teacher supervision al...
Jun-Chi Chen, Chang-Tao Miao, Yu Xiang et al.· 0 citations
This work presents a novel approach to monitoring the network traffic of agentic systems using complex valued hypersparse traffic matrices by integrating DBOS (DataBase OS), the OneSparse PostgreSQL database, and the GraphBLAS math library.
Manuel Tsoukatos, Hayden Jananthan, Jeremy Kepner· 0 citations
Auditing publicly released descendants of distilled models spanning diverse post-training objectives, SCOUT consistently identifies the distillation source and tracing teacher-associated *syntactic signatures* along training trajectories reveals that they emerge during distillation and persist through subsequent prefer...
Minwoo Jang, Jaechang Kim, Minhyeon Oh et al.· 0 citations
Results show that preserving already-correct work under unsupported accusation is a distinct safety challenge for long-lived agents, and introduces CAVE-Bench, a benchmark of 365 agentic tasks across six domains built around opaque tasks.
Xu-Tao Mao, Rui Qian, Long-Xiang Wang et al.· 0 citations
A probabilistic risk model linking five stages: reward hacking, containment escape, usable access, persistence, and failure of detection supports treating cyber-capable agent evaluations as hostile security zones in which indirect egress, shared infrastructure, credentials, and evaluation artifacts must remain outside...
Murat Ozer, Bulent Erenay, Ibrahim Berber· 0 citations