Skip to content

Category

cybersecurity

1,065 papers

#artificial intelligence Preprint Sep 2026

Where Cyber Agents Struggle: Bottleneck Analysis of Multi-Stage LLM Agents

This work presents an end-to-end diagnostic study of an Autonomous Adversary system with orchestrator, executor, and validator LLMs in enterprise-like lateral-movement scenarios and uses comparative LLM-as-a-Judge analysis to identify planning deficiencies.

Saeedeh Lohrasbi, Mohammad Mamun, Ahmed Yehia et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Who Is Behind the Harness? Fingerprinting LLMs through Agentic Behavior

LLMs increasingly operate through coding-agent harnesses that inspect repositories, invoke tools, and modify files. Substituting the model behind such an agent can therefore change security-relevant decisions, including whether it verifies changes or recovers safely from failures. Existing LLM fingerprints largely infe...

Chu-Yi Wang, Xiao-Hui Xie, Tong-Ze Wang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

PrivDrift: Auditing User-Secret Leakage Under Topic Drift in Active LLM Conversations

PrivDrift is introduced, a benchmark for auditing whether user-disclosed secrets remain recoverable after conversational topic drift and persuasion-based probing, suggesting that privacy risk in active LLM contexts should be evaluated as a persistent behavioral failure mode rather than only as training-data memorizatio...

L. Maldonado · 0 citations
#artificial intelligence Preprint Open access Sep 2026

ENDOPROMPT: Victim-Side Pseudo-References for Utility Degradation

Prompt injection can degrade benign task performance without eliciting harmful content. Yet many attack objectives depend on task labels or predefined target responses. We present ENDOPROMPT, a white-box method that learns utility-degrading prefixes from unlabeled instructions. Its generator takes the request text as i...

Qingyu Wu, Zeyu Feng, Yongda Yu et al. · 0 citations
#artificial intelligence Preprint Aug 2026

PartHackBench: Certified Equal-Progress Stress Tests for Partial-Credit Tool-Agent Evaluation

PartHackBench provides a certified control for testing whether evaluator credit changes while all benchmark-defined task-relevant progress remains fixed, and provides a certified control for testing whether evaluator credit changes while all benchmark-defined task-relevant progress remains fixed.

Hong-Ye Yang, Zhi Xie, Sheng-Jun Xiong · 0 citations
#artificial intelligence Preprint Open access Sep 2026

Just Ask Jev: Reinforcement Learning for Calibrated Decisions as a Zero-Shot Detector of AI Alignment Failures

Detectors of alignment failures screen deployed language models and score alignment benchmarks. Most are generative judges that spend a decoding pass on every criterion, and classifiers that read token probabilities, such as Llama Guard, still score one fixed label per call. Jev, a model trained with reinforcement lear...

Ruoqi Guo, Yi Liu, Gelei Deng et al. · 0 citations
#artificial intelligence Preprint Open access Sep 2026

Control the Harness, Control the Cost: Routing and Governing AI Coding Agents in the Enterprise

Harnesses, the products that run AI coding agents, are multiplying, and enterprises are rolling them out to their employees: what started as pilots with a few hundred seats is scaling to tens of thousands. Most enterprises do not build these harnesses but buy them from large vendors, such as Anthropic's Claude Code or...

Arian Abbasi, Alan Aqrawi, Ted Kwartler · 0 citations
#artificial intelligence Preprint Sep 2026

Progressive Skill Discovery as Access Control for Tool-Using LLM Agents: Structural Governance through Role-Scoped Capability Delivery

Skilder is introduced, a framework that packages capabilities into roles: bundles of skills, tools, and instructions, together with the limits that bound them, and preserves problem-solving flexibility while providing hard system-level enforcement.

Michael Stettler, Benjamin Girardet, Jonas Canton et al. · 0 citations
#artificial intelligence Preprint Aug 2026

A Non-Invasive Cloud-Based Migration Strategy for Post-Quantum Cybersecurity in Smart HVAC Systems: Architecture, Implementation, and Empirical Evaluation

Legacy smart HVAC controllers rely on vendor-cloud TLS secured by ECDH and RSA, both broken by Shor's algorithm, and typical 10-15 year lifespans mean today's devices remain in service through the quantum-threat era. Direct on-device post-quantum cryptography is infeasible: an ESP32-S3, representative of capable HVAC h...

Mahedee Zaman Moon, Kaysarul Anas Apurba, Mahamudul Hasan et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Issuer-Sovereign Agentic Payments

AI agents are beginning to make real payments. Current approaches let an agent pay by relying on a credential provider that, in the approaches deployed today, typically sits outside the cardholder's bank. The spending rules are then enforced by the card network or that provider, and not by the bank itself. This leaves...

Disha Sharma, R. Kaushal, Ashu Kanaujia · 0 citations
#artificial intelligence Preprint Sep 2026

Multi-View Fusion for Encrypted C2 Detection: A Leakage-Controlled Measurement Study of Evaluation Pitfalls

This work reports three findings that matter more than the fusion result itself, and concludes that for encrypted C2 detection, the evaluation design is not a preliminary step and fusion beats the best single view by only 0.022 in F1.

H. Nguyen-Huu, V. Phan, Khuong Nguyen-An · 0 citations
#artificial intelligence Preprint Sep 2026

When Clients Are Orchestrated: Strategic Gradient Manipulation to Defeat Federated Learning Servers with Efficient Defense

Federated Learning enables decentralized model training by exchanging model updates--rather than raw data--with a central parameter server (PS). While most of the existing defenses primarily assume static or independently acting adversaries, we reveal a new class of dynamically adaptive attacks that systematically bypa...

M. Shaaban, A. Abdel-Naby, Mohamed Elmahallawy · 0 citations

From tech blogs

See all →
Google DeepMind Blog Jul 17, 2026

Introducing Gemini 3.5 Flash Cyber

Google introduces Gemini 3.5 Flash Cyber, a lightweight cybersecurity model to find and patch vulnerabilities.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.