Dagger, a novel two-phase decoupling-based attack framework that consistently outperforms state-of-the-art GNN stealing attacks, achieving up to 18.16\% higher fidelity while only utilizing 12.23$\times$ fewer queries than the strongest baseline.
This paper presents the design, formal analysis, and implementation of a complete blockchain-based electronic invoice system on Ethereum, and formalizes the invoice lifecycle as a guarded labeled transition system and proves that the system guarantees reimbursement uniqueness, face integrity, and authorization soundnes...
Driven by the rapid advancement of large language models (LLMs), LLM-based multi-agent systems (MAS) have emerged as a powerful paradigm for collaborative reasoning over complex tasks. A key design element of MAS is the communication topology, which governs information flow among agents and often encodes proprietary kn...
Long-Zhu He, Ze-Kun Wen, Xin-Feng Li et al.· 0 citations
This work proposes a controlled inject-and-remove cycle: deliberately inject a weaker backdoor and then unlearn it, which weakens detector-visible signatures and fools the backdoor detectors with an illusion of purification while preserving the malicious retrieval behavior.
The proposed HammingMark is a semantic watermarking method that uses the semantic hash of the preceding sentence as a dynamic center and accepts candidates whose hashes fall within its Hamming neighborhood, which achieves strong robustness, high detectability, and near-unwatermarked generation quality.
Ze-Wen Sun, Tong-Yang Zhao, Li-Yao Xiang et al.· 0 citations
This work introduces ToolFence, which compiles a typed authorization blueprint before execution, enforces it through a deterministic monitor, and when the blueprint is incomplete asks a judge to grant new capabilities rather than adjudicate each concrete call, improving runtime efficiency.
Yan-Jie Li, Xiang-Yu He, Xue-Long Dai et al.· 0 citations
OPFL calibrates an empirical boundary offline and uses it to distinguish benign numerical deviations from malicious manipulation, and adopts optimistic verification by post auditing only sampled training steps to reduce the cost of expensive MPC replay.
Hong-Xu Su, Jian-Zhu Yao, Xue-Chao Wang et al.· 0 citations
Repeated verifier calls are useful only when they contribute conditional information. We introduce VStress, an auditable replay contract, and VStress-CA, a correlation-aware allocation policy that estimates the conditional marginal information of an unqueried verifier on a sealed calibration split, discounts uncertaint...
Miao-Bo Hu, Shu-Hao Hu, Xiao-Bo Guo et al.· 0 citations
This work introduces \method{}, a framework for jailbreaking through text-only continuation interfaces that permit repeated sampling and assistant-prefix continuation, and achieves the highest mean score most comparisons against baselines.
Jesson Wang, Shawn Li, Wei Yang et al.· 0 citations
Experiments show that SKILLLITE improves malicious Skill detection across different compact LLM backbones and outperforms existing representative auditing baselines, and generalizes to behaviorally confirmed in-the-wild malicious Skills.
Hao-Ran Ou, Ge-Lei Deng, Xuan-Ye Zhang et al.· 0 citations
Fine-tuning-as-a-service lets users adapt a safety-aligned language model to their own data, but it also creates a harmful fine-tuning attack surface: a small amount of harmful data mixed into an otherwise benign fine-tuning set can degrade the model's alignment. Two recent alignment-stage defenses address this problem...
Muhammad Zeeshan Akram, Mufid Kamel Marican, Anvesh Reddy Yenugu et al.· 0 citations
Safety-aligned language models are commonly deployed as multi-turn assistants, which lets adversaries spread unsafe intent across several user turns instead of a single prompt. Gradient-based jailbreak detectors such as GradSafe were developed for single prompts: they score an input by the alignment between its induced...
Omar Sheta, Rinku Deuja, Hadi Masoudi et al.· 0 citations