Skip to content

Author

Wei Zhao

We have 3 of 16 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#artificial intelligence Preprint Sep 2026

Representation Transitions Reveal Emerging Safety Risks in Multi-Turn LLM Agents

Multi-turn attacks on agentic systems can compose individually permissible actions into harmful outcomes, challenging defenses that assess actions or states in isolation. We show that such attacks leave a detectable signature in the agent's internal representations: harmful behavior emerges as an accumulated representa...

Hao-Yu Wang, Wei Zhao, Ye-Di Zhang et al. · 0 citations
Preprint Aug 2026

Tripwire: Triggering Aligned Refusal via Statistically Certified Safety Neurons

Neuron- and path-level interventions offer the finest-grained route to defending large language models (LLMs) against jailbreak attacks, yet existing methods fall short of this promise, i.e., they often compromise model utility significantly. Specifically, one line of work suppresses toxic neurons to erase harmful sema...

Wei Zhao, Zhe Li, Pei-Xin Zhang et al. · 0 citations
Preprint Jun 2026

Yuvion LLM: An Adversarially-Aware Large Language Model for Content And AI Safety

As large language models are increasingly deployed in real-world systems, safety failures can still lead to harmful outputs and dangerous misuse. We argue that the essence of safety is adversarial: many failures arise not from natural inputs alone, but from strategic attempts to evade model policies and safeguards. How...

Ting Ma, Xiufeng Huang, Benlei Cui et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.