Skip to content

Author

Hongtao Wang

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#artificial intelligence Preprint Oct 2026

SLDR: Defending Against Malicious Fine-tuning via Selective Layers Recovery and Dynamic Routing

Fine-tuning-as-a-service enables users to adapt aligned large language models (LLMs) to specialized tasks, but malicious fine-tuning can erode refusal behavior while preserving task performance on legitimate inputs. We revisit recent layer-wise safety diagnostics and find that safety sensitivity is signed: scaling diff...

Hui Zhang, Ya-Chao Yuan, Jia-Yun Wang et al. · 0 citations
#artificial intelligence Preprint May 2026

MemPoison: Bypassing Selective Memory Mechanisms to Plant Backdoors in LLM Agents

MemPoison is proposed, a novel memory poisoning attack that bypasses selective memory mechanisms in LLM agents, where an attacker can inject triggerable backdoors into the agent's long-term memory through dialogue interactions, thereby misleading its subsequent responses.

Hongtao Wang, Sean Yang, Yu Chen et al. · 5 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.