Skip to content

Category

cybersecurity

1,065 papers

#artificial intelligence Preprint Sep 2026

High-Capacity Robust Medical Image Exfiltration via Neural Network Weight Replacement

Collaborative medical AI platforms allow researchers to train models on sensitive imaging data while restricting data export. However, trained models can serve as covert carriers of patient information: medical images may be encoded within model parameters and reconstructed outside the secure environment. Existing defe...

Elie Thellier, Hui-Yu Li, Nicholas Ayache et al. · 0 citations
#artificial intelligence Preprint Open access Sep 2026

What Drives Dialectal Jailbreaks? An Ablation of Surface Form, Cultural Framing, and Strategy Banks

Recent work suggests that obscure language registers can weaken large language model refusal behavior, especially when paired with black-box prompt optimization. It remains unclear whether failures stem from non-standard surface form, culturally grounded framing, or optimization over an expressive prompt-strategy space...

Qingyang Xu · 0 citations
#artificial intelligence Preprint Sep 2026

Information Design Against Gaming and Learning Adversaries

The Pareto frontier between the two defense objectives is characterized, and both rates are confirmed on seven binary-classification tasks spanning tabular, image, and language-model-feature inputs: label-plus-counterfactual access extracts the boundary with up to $200\times$ fewer queries than a published label-only b...

Madhava Gaikwad · 0 citations
#artificial intelligence Preprint Sep 2026

Share-Borne AI Virus: Memory-Hopping Attacks Across LLM Agents

This work introduces artifact-mediated propagation, where adversarial content introduced through an artifact is stored in an assistant's persistent memory, reproduced in a subsequently created artifact, and acquired by another assistant that later reads it.

Sidharth Pulipaka, Anshu Sharma, Stanislau Hlebik et al. · 0 citations
#artificial intelligence Review Sep 2026

When Valid Tool Calls Change Meaning: Formation-Consistent Dispatch for LLM Agents

Formation-consistent dispatch (FCD), which connects implementation analysis to execution authority, is presented, which produces provenance-bound over-approximations of declared in-scope effects from official source.

Geonwoo Kim, Brent ByungHoon Kang Korea Advanced Institute of Science, Technology · 0 citations
#artificial intelligence Preprint Sep 2026

Before Acting, Change the State: Prospective State Intervention for Web Agents under Deceptive Interfaces

Veer is introduced, an agent-side runtime defense that leaves task planning to the base agent and intervenes on Web state when a proposed action would produce an unauthorized consequence, establishing task-relevant Web state as an effective runtime control target for protecting Web agents from deceptive outcomes.

Ruo-Zhao Yang, Ming-Fei Cheng, Xiao-Fei Xie · 0 citations
#artificial intelligence Preprint Sep 2026

PlanGuard: A Guardrail for Multi-Step Plan Safety in Embodied Agents

This work introduces PlanGuard, the first pre-execution detector that evaluates the physical safety of a complete multi-step plan in its current environment, and proposes Strong-Teacher Adaptive Compensation for On-Policy Distillation (STAC-OPD), which provides compact models with adaptive strong-teacher supervision al...

Jun-Chi Chen, Chang-Tao Miao, Yu Xiang et al. · 0 citations
#artificial intelligence Review Sep 2026

Agentic Network Traffic Monitoring

This work presents a novel approach to monitoring the network traffic of agentic systems using complex valued hypersparse traffic matrices by integrating DBOS (DataBase OS), the OneSparse PostgreSQL database, and the GraphBLAS math library.

Manuel Tsoukatos, Hayden Jananthan, Jeremy Kepner · 0 citations
#artificial intelligence Preprint Sep 2026

Retrospective Distillation Attribution via Normalized Response Similarity

Auditing publicly released descendants of distilled models spanning diverse post-training objectives, SCOUT consistently identifies the distillation source and tracing teacher-associated *syntactic signatures* along training trajectories reveals that they emerge during distillation and persist through subsequent prefer...

Minwoo Jang, Jaechang Kim, Minhyeon Oh et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Reward Hacking and Agent Containment Failure: A Monte Carlo Study Based on the 2026 Hugging Face Incident

A probabilistic risk model linking five stages: reward hacking, containment escape, usable access, persistence, and failure of detection supports treating cyber-capable agent evaluations as hostile security zones in which indirect egress, shared infrastructure, credentials, and evaluation artifacts must remain outside...

Murat Ozer, Bulent Erenay, Ibrahim Berber · 0 citations

From tech blogs

See all →
Google DeepMind Blog Jul 17, 2026

Introducing Gemini 3.5 Flash Cyber

Google introduces Gemini 3.5 Flash Cyber, a lightweight cybersecurity model to find and patch vulnerabilities.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.