This work compares a scaled-Manhattan matcher, a gradient-boosted classifier, a gradient-boosted classifier, a TypeNet-style recurrent embedding model, and a TypeFormer-style Transformer under a 5-fold subject-disjoint protocol and a design that jointly varies mechanism and the enrolment-to-query gap, and proves large...
Siôn Parkinson, Saad Khan, Na Liu et al.· 0 citations
This monograph presents a first-principles forensic autopsy of the intrusion, provides formal evidence that the breach was a predicted consequence under the Instrumental Convergence thesis operating within an unattenuated autonomous loop lacking out-of-band circuit-breakers, exposes the Defensive LLM Guardrail Paradox...
This is the first systematic, controlled study that isolates the scratchpad reasoning channel as an output-prefix attack vector, and the first to compare reasoning-only, output-prefix-only and reasoning-plus-output-prefix attacks across both exposed- and hidden-reasoning models.
Lukáš Brůna, Robert A. Bridges, Adam Ek· 0 citations
It is shown that feedback-based agents can retain early biases even when later correction is available, and a black-box framework for exploiting this weakness is proposed that operationalizes the three factors as directional-shift, contextual-plausibility, and counterevidence-resilience signals under either limited tar...
Chuan-Chao Zang, Jia-Ning Wang, Wen-Yu Chen et al.· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
This work argues that agents need an operating-system substrate providing mandatory, non-bypassable services for identity, input mediation, memory governance, and execution control, and introduces AgentKernel, a trust-native agent operating system built around the premise that security must be a first-class design cons...
Zhen-Hua Zou, Sheng Guo, Qiu-Yang Zhan et al.· 0 citations
TP-CRIV targets a third-party verification setting in which the verifier has neither white-box nor API access to the claimant's model, can interact with the suspicious deployed service only through its ordinary black-box inference interface, and does not require protocol-specific cooperation from the service provider.
Toketive is introduced, a simple yet powerful reference-free attack that exploits the tokenization-based side channel to detect modified knowledge and reconstruct the corresponding pre-edit response and shows that localized modifications should not be treated as robust knowledge-control boundaries without adversarial e...
Autonomous penetration-testing harnesses use large language models (LLMs) for reconnaissance, exploitation, and reporting, but often rely on those same models to confirm findings, grade severity, and select agents. This can lead to false positives, inflated severity, and wasted compute. We examine how System One decisi...
This work pairs application-level agent telemetry with kernel-level syscall traces to present the first paired-evidence characterization of kernel-level versus application-layer signal for agent security, finding that kernel evidence is discriminative on its own and that composing it with application-layer evidence gen...
Spencer King, Zhi-Lu Zhang, Mikhail Kuznetsov et al.· 0 citations
Blockchain and artificial intelligence (AI) are converging into a single infrastructural layer for securing data sharing, model integrity, and autonomous decision-making across distributed systems. This paper presents a meta-synthesis that draws together four constituent studies covering adversarial machine learning, A...
It is shown that schema-defined outputs change but do not eliminate prompt-injection risk, highlighting the need to evaluate how untrusted content influences choices within the allowed action set.
Multi-step tool-calling LLM agents rely on host runtimes to preserve state across turns. When a runtime carries an external tool return into later model inputs, providers meter it again. An admitted malicious or compromised tool can thereby convert untrusted data into recurring victim-billed processing without victim c...
Jin-Qian Zhang, Hao-Jun Xia, Shu-Jiang Wu et al.· 0 citations