Skip to content

Category

cybersecurity

1,065 papers

#artificial intelligence Preprint Sep 2026

Retrosynthesis of Synthetic Media for Explainable AI Provenance Forensics

A self-referential retrosynthesis framework for explainable AI provenance forensics under a fixed-generator setting that leverages a jointly optimized encoder-decoder pair to implement a self-embedding mechanism that enables round-trip consistency verification.

Yi-Jie Lin, Ching-Chun Chang, Isao Echizen et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Transfer Safety Awareness for Cross-Modal Safety Drift in Multimodal Large Language Models

Motivated by the observation that safety signals from unsafe text processing can be transferred, safety-awareness representation transfer (SRT) is proposed, a lightweight direction-refinement method that mitigates cross-modal safety drift with a frozen MLLM backbone.

Tian-Qi Xiao, Shiyao Cui, Ming-Hao Zhang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Agent Memory Is a Surface for Endogenous Authorization Laundering

This work evaluates five LLMs as memory writers and two as executors across procurement, cybersecurity, and finance and introduces EAL-Bench, which measures how accurately persistent memory preserves evolving authorization state and whether errors propagate to downstream unauthorized actions.

Tommaso Cerruti, Mika Okamoto, Ansel Kaplan Erol · 4 citations · ⚡1
#artificial intelligence Preprint Open access Sep 2026

Public-Sharing Labels and Verbatim Field Egress in an MCP-to-A2A Agent Configuration: A Controlled Multi-Model Study

Safety properties assessed separately for Model Context Protocol (MCP) tool use and Agent2Agent (A2A) delegation need not describe behavior when one agent uses both. We measure one such behavior in a single controlled MCP-to-A2A configuration: a testbed drives a real-model host across a local MCP and a local A2A leg in...

Arpan Kumar Mahapatra · 0 citations
#artificial intelligence Review Sep 2026

Ranked by the Matcher: A Reproducibility Audit of Knowledge Graph Extraction from Threat Reports

To separate component effects from matcher rewards, CTIForge is built, whose deterministic validation layer can vary while extraction is held byte-identical, and which coincides with a roughly 2.8-fold increase in actions explicitly disputing entity type.

Safayat Bin Hakim, H. Song · 0 citations
#artificial intelligence Preprint Sep 2026

SafeEvolve: Harness-Policy Co-Evolution from Agent Experience for Safety Alignment

Through harness-policy co-evolution, SafeEvolve converts safety experience into an evolved runtime harness and improved policy behavior, and experiments show that SafeEvolve achieves a stronger safety-utility tradeoff than existing baselines.

Qing-Hua Mao, Wanying Qu, Da-Di Guo et al. · 6 citations · ⚡2
#cybersecurity Review Sep 2026

SoK: Motion Data Privacy in Extended Reality

This SoK examines 134 relevant papers on privacy concerns in motion patterns recorded by XR headsets, including how adversaries can obtain users'motion patterns, the inferences they can draw from them, and methods for protecting users and clarifies the state of XR motion privacy.

Azim Ibragimov, Alina Vasina, Uliana Polshcha et al. · 1 citation
#natural language process... Preprint Feb 2025

GuidedBench: Measuring and Mitigating the Evaluation Discrepancies of In-the-wild LLM Jailbreak Methods

GuidedBench, a novel benchmark comprising a curated harmful question dataset and GuidedEval, an evaluation system integrated with detailed case-by-case evaluation guidelines are introduced, ensuring reliable and reproducible evaluations.

Ruixuan Huang, Xun-Guang Wang, Zongjie Li et al. · 8 citations
#natural language process... Preprint Sep 2026

Membership Inference in Fine-tuned Diffusion Language Models via Token-level Memorization Asymmetry

This work identifies a phenomenon termed token-level memorization asymmetry through theoretical analysis of diffusion training dynamics and proposes Q-Skew, a quantile-weighted skewness-based indicator for membership inference on finetuned DLMs.

Sheng-Fang Zhai, Leo Marchyok, Yu-Ling Shi et al. · 1 citation · ⚡1
#natural language process... Preprint Aug 2026

TRIS: A Tri-Layer Retrieval Integrity Sieve Against Knowledge Poisoning

The Tri-Layer Sieve is presented, a middleware defense that sanitizes retrieved evidence through cross-embedding-space clustering with an independent judge model, structural filtering of trigger-payload artifacts, and LLM consistency verification, and exploits a key weakness of retrieval-stage poisoning.

Muhaimin Bin Munir, Akib Jawad Ononto, Nazia Shehnaz Joynab et al. · 0 citations

From tech blogs

See all →
Google DeepMind Blog Jul 17, 2026

Introducing Gemini 3.5 Flash Cyber

Google introduces Gemini 3.5 Flash Cyber, a lightweight cybersecurity model to find and patch vulnerabilities.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.