This work shows that, despite disrupting the spatial structure required by conventional reconstruction attacks, transmitted token embeddings retain substantial positional information, and introduces the Spatially Aligned Reconstruction Attack (SARA), a unified pipeline that predicts token positions, restores their spat...
C-SafeQA, a policy-grounded benchmark for response-level Chinese safety evaluation, is introduced, with substantial trade-offs between unsafe-response recall and risk-query-conditioned safe-response false positive rate.
Rui-Wei-Chun-Hui-Li-Yuan Yang, Shuang Huang, Jun-Hua Liu et al.· 0 citations
Athea, the first graph-based approach for vulnerability affected library identification, is proposed, which models vulnerability databases as a knowledge graph and reformulates the identification problem as knowledge graph completion (KGC).
P. Duy, Trang Dang Yen, Hưng Nguyễn Hữu et al.· 0 citations
This work presents HiveTraceGuard-Pro, a 0.6B generative guardrail LoRA-tuned from Qwen3-0.6B that has the highest clean Russian robustness combined-F1 and Russian prompt-injection recall and releases the merged weights on Hugging Face under Apache-2.0.
N. Oblakov, Sabrina A. Sadiekh, Evgeniy Kokuykin· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
DP-TabImage is proposed, a modality-specialized framework for private paired synthesis that achieves a strong balance among tabular fidelity, image fidelity, and cross-modal alignment and reveals that visual warm-up primarily improves marginal image fidelity, whereas aligned table-image warm-up is critical for improvin...
Kai Chen, Josephine Lamp, S. Jha et al.· 0 citations
This work systematizes MAS security through an execution-centered analysis of 197 works, introducing an A-I-R framework that organizes attacks by adversary position, interaction interface, and resulting system-level risk, unifying otherwise fragmented attack mechanisms across MAS.
Rui Yang, Jun-Jie Xu, Zheng-Yu Liu et al.· 1 citation
This work determines what each reported result implies for that question: how much help with harmful tasks the service still gives an attacker who keeps adapting or finds another way in.
This work argues that red-teaming is better framed as a search problem: discover, organize, and iteratively refine a diverse archive of attack strategies, producing a structured map of how a target model fails rather than a list of one-off successes.
Fei-Tong Qiao, Li-Ren Peng, Shi-Ming Ren et al.· 0 citations
This paper claims that an in-depth investigation of the adversarial robustness of NeSy models is necessary and provides the first systematic evaluation of backdoor attacks against NeSy, and shows that while NeSy models are indeed more robust than their neural counterpart on average, their robustness vastly depend on th...
Marco Antonio Corallo, Andrea Agiollo, Mauro Conti et al.· 0 citations
Deployed language model safeguards (safety fine-tuning, filtering, unlearning) vary by principal only outside the model weights: filters are reconfigured, tiers are multiplied, and artefacts are reissued; inside one set of weights every request meets the same model configuration. This motivates us to define capability-...
Patrikas Vanagas, Augustas Macijauskas, L. Lopata· 0 citations
It is shown that an external observer can identify the class of the workload running on an NVIDIA H200 from its power draw, and adversarial traces are recorded, offering initial insights beyond genuine activities and a dataset for developing and testing stronger evasion mechanisms.
It is argued that agent security must be evaluated under an untrusted-model assumption: a correct system is one in which a fully prompt-injected agent still cannot exceed the authority explicitly delegated to it, and an authorization broker is implemented that closes the gap.
Panduranga Sai Varma Dantuluri, Jyotirmoy Sundi· 0 citations