NeuroBreak is presented, a visual analytics system that helps experts progressively unpack jailbreak mechanisms from layer-level semantics down to neuron-level behaviors and provides actionable insights for strengthening LLM defenses.
Chuhan Zhang, Ye Zhang, Bowen Shi et al.· arXiv.org· 3 citations
It is proved that, under an intuitive and experimentally supported assumption called distribution consistency, obfuscation can nullify the robustness of N-gram-based watermarks and motivate more semantics-aware alternatives.
Gehao Zhang, Eugene Bagdasarian, Juan Zhai et al.· arXiv.org· 0 citations
The first Fraud Investigation Assistant (FIA) framework is introduced, which employs multimodal large language models (LLMs) to automate key steps of credit card fraud investigation and generate explanatory reports and suggests that LLM-based agents can assist with automating substantial parts of the fraud investigatio...
Shaun Shuster, Eyal Zaloof, A. Shabtai et al.· arXiv.org· 5 citations· ⚡1
REP, a lightweight in-context elicitation method that uses shadow-model-generated demonstrations wrapped in auxiliary code-like formats to raise user-visible reasoning traces from a victim model, substantially increases similarity between exposed and REP-conditioned internal traces while preserving useful reasoning sig...
A four-stage forensic audit protocol for API-served models, validated prospectively on a flagship case whose 2026-08-23 analysis pointed to the GLM-5.3 version line and whose official reveal confirmed those family and version-line inferences.
SIR is presented, a black box IPI attack that composes stealthy injections from a small library of reusable principles stated in plain language and wraps composition in an iterative feedback loop that diagnoses the victim's failed trajectories and distills the bypasses into new, named strategies that are reapplied acro...
Chen Xiong, Zhi-Yuan He, Pin-Yu Chen et al.· 0 citations
The results show that the studied causal signal reveals what shaped an action without reliably encoding whether the action was authorized, and that reference construction and routing are integral to the effective security decision.
Tanzim Ahad, Ismail Hossain, Md. Jahangir Alam et al.· 0 citations
This paper presents the first systematic security study of checkpoint and rollback in existing agent systems, and develops a multi-agent analysis pipeline that reconstructs execution semantics, identifies violations of the five failure conditions, and validates them through actual rollback.
Guan-Long Wu, Da-Hui Li, Ke Jiang et al.· 2 citations
Results show that protective intervention is sensitive to the surface through which a request arrives, that this sensitivity is detectable using a simple protective coding scheme, and that it is not explained by turn length alone.
E. Lee, Soo-Young Lee, Jungpyo Nam et al.· 0 citations
CoVeil is proposed, a defense mechanism which dynamically optimizes transmitted signals to suppress leakage during decoding time while preserving the collaborative quality, and consistently improves the privacy-utility trade-off over existing baselines by reducing data leakage.
Ke-Jia Zhang, Tianyuan Zou, Zi-Xuan Gu et al.· 0 citations
This work evaluates escalation channels, structured reporting tools available to the agent at the point of conflict, as a decision-environment intervention that both reduces reward hacking and surfaces the infrastructure defects that trigger it.
The findings suggest that ADBA offers solid advantages to the traditional PINs, and successfully addresses the tradeoffs between efficiency, memorability, and security under the usage scenarios considered in the study.
Yuxuan Huang, Qiao Jin, Tongyu Nie et al.· 0 citations