Cybersecurity incidents such as data breaches have become increasingly common, affecting millions of users and organizations worldwide. The complexity of cybersecurity threats challenges the effectiveness of existing security communication strategies. Through a systematic review of over 3,400 papers, we identify specif...
Carolina Carreira, Alexandra Mendes, Jo\~ao F. Ferreira et al.· 0 citations
A dual-assessment framework is introduced that scores any real-time gaze transformation on two axes at once: interaction utility, measured through an offline gaze-interaction simulation and an operational spatial-accuracy metric, and privacy preservation, measured through the Rank-1 Identification Rate of a state-of-th...
Mehedi Hasan Raju, Oleg V. Komogortsev· 3 citations
Across model-level safety on HarmBench and agent-level safety on AgentHarm, Membrane achieves the highest F1 on all six modern jailbreak attacks, and benign refusal on AgentHarm stays at 7-14%, well below the 28-85% range of prior guards.
Minseok Choi, Seung-Moo Yang, Dongjin Kim et al.· arXiv.org· 2 citations
We reveal that Deep Research (DR) agents systematically expose safety risks: simply submitting harmful queries that a standalone LLM would reject outright can elicit detailed and dangerous reports from DR agents. Empirical analysis reveals that the advantages that make DR agents powerful unintentionally make them vulne...
Shuo Chen, Zonggen Li, Xingyu Jin et al.· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
Open-weight Large Reasoning Models (LRMs) are approaching the capabilities of their frontier counterparts but pose significant safety concerns, as they are difficult to patch or monitor post-release. To prevent misuse, reasoning-based safety guardrails, where models explicitly reason on safety justifications before ans...
Shuo Chen, Zhen Han, Haokun Chen et al.· 0 citations
This work proposes two classes of attacks: a static attack, which employs a fixed injection prompt, and an iterative attack, which optimizes the injection prompt against a simulated reviewer model to maximize its effectiveness.
Qing Zhou, Zhexin Zhang, Zhi Li et al.· arXiv.org· 6 citations· ⚡1
With the widespread applications of large language models (LLMs), privacy-preserving inference has become increasingly essential for sensitive queries. To balance privacy and utility, a series of lightweight obfuscation approaches has recently been proposed, where users locally transform plaintext embeddings into the f...
Si-Cong Li, Ling-Feng Yao, Xing-Ke Yang et al.· 0 citations
Aligned language models fail under two independent pressures: the structural jailbreak class recently formalized as Involuntary In-Context Learning (IICL), which reframes a harmful request as the final missing cell of a data-labeling task completed by pattern rather than judged as content; and the erosion of safety ali...
A privacy-preserving federated learning framework for clinical EEG data that uses masking-based secure aggregation as its core protection mechanism that remains compatible with federated model training, although malicious-setting safeguards and lightweight consistency-checking mechanisms introduce additional computatio...
Recent cryptographic results establish that neural networks can be backdoored such that no efficient algorithm can distinguish them from a clean model. These guarantees, however, have been confined to stylised architectures of limited practical relevance, leaving open whether comparable undetectability extends to moder...
Marte Eggen, Eirik Reiestad, Kristian Gj{\o}steen et al.· 0 citations
Machine learning systems deployed in distributed or federated environments are highly susceptible to adversarial manipulations, particularly availability attacks -- rendering the trained model unavailable. Prior research in distributed ML has demonstrated such adversarial effects through the injection of gradients or d...
We present Tempora-Fusion, the first homomorphic TLP scheme with efficient public verification of both individual puzzle solutions and homomorphic linear combinations. Tempora-Fusion lets clients generate puzzles independently, later authorize a linear combination with its own release time, and enables any party to ver...