A federated learning framework that combines a Gaussian-mechanism differential privacy layer with a coordinate-wise trimmed-mean Byzantine-robust aggregation rule, evaluated on a simulated cross-institutional classification task resembling fraud and clinical-risk scoring.
A deployed system that scores an address by its position in a multi-chain transaction graph rather than by its presence in a list, and an adversarial harness of eight recurrent reinforcement-learned archetypes that passes an 8-criterion degeneracy audit and exposes a measured blind spot of the deployed heads against sy...
As vehicular networks move toward 5G/6G edge intelligence, federated learning (FL) is widely promoted as a privacy-preserving way for vehicles and infrastructure to train shared models without exposing raw sensor data. Yet the updates clients transmit still leak enough information to identify who sent them, which threa...
Ali Akarma, Toqeer Ali Syed, Muhammad Khan et al.· 0 citations
This work studies refusal directions through the training dynamics across refusal datasets and reveals that their brittleness is associated with repetitive refusal starts, which is linked to concentration of gradients and refusal features in a low-dimensional subspace.
Andrey Labunets· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
From pre-training to query-time augmentation, web-scraped data helps to improve the quality and contextual relevancy of content generated by large language models (LLMs). However, large-scale web scraping to feed LLMs can affect site stability and raise legal, privacy, or ethics concerns. If website owners wish to limi...
Steven Seiden, Triss Ren, Caroline Zhang et al.· 0 citations
The results show that context-aware security assessment is a practical complement to existing AE processes and can support safer and more responsible research artifact sharing.
The Model Context Protocol (MCP) widens the prompt injection attack surface of large language model applications to tool descriptions, parameter schemas, and tool outputs. Defenses for it are appearing quickly, but their reported figures are not comparable: each is evaluated on a corpus of its authors' construction, un...
\.Ipek Abas{\i}kele\c{s} Turgut, Edip G\"um\"u\c{s}· 0 citations
A new benchmark to more comprehensively evaluate CUAs'misuse risks, CUAHarm, and explores using LMs to monitor CUAs'actions, finding monitoring unsafe computer-using actions is significantly harder than monitoring conventional unsafe chatbot responses.
In this paper, we investigate whether refusal behavior can be predicted from LLM intermediate activations before decoding using linear probes trained on residual stream activations at each transformer block. We find that refusal is linearly decodable well before the final layer, indicating that safety-relevant behavior...
Matteo Gioele Collu, Riccardo Conte, Alberto Giaretta et al.· 0 citations
This work presents Sentinel-RL, an agentic-SOC architecture that decouples topological reasoning from semantic reasoning, and contributes a reusable engineering pattern, a portable HPC deployment pattern, and an enterprise-readiness analysis covering false-positive economics, reversibility guarantees, audit compliance,...
Uday Vallabhaneni, Cassie L. Cagwin, David J. Wild· 1 citation
For every coherent and sufficiently expressive finite syntactic system S, we prove the existence of at least one theorem that S cannot produce autonomously. The result is a metatheorem: it proves the existence of a theorem, and applies to every finite syntactic system - security mechanisms, AI systems, formal verifiers...
A patch similarity metric is introduced to detect memorized patches and new patch validation methods are developed that thoroughly evaluate both security and semantic correctness of agent patches for vulnerability patching.
Chihao Shen, Jia-Cheng Li, Aastha Mahajan et al.· 0 citations