Skip to content

Author

Luo-Yu Chen

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#artificial intelligence Preprint Oct 2026

Target-free Latent Safety Alignment

Large language models (LLMs) remain highly vulnerable to jailbreak attacks that induce harmful behaviors and circumvent safety alignment. To defend against such attacks, adversarial training paradigms have been proposed to first simulate failure modes and then train the model to correct them, yielding promising improve...

Luo-Yu Chen, Wei-Qi Wang, Chen-Han Zhang et al. · 0 citations
#artificial intelligence Preprint Oct 2026

Reactivating Alignment: Defending LLMs from Jailbreaks via Intention-Aware Input-Output Matching

Large language models (LLMs) remain vulnerable to jailbreak attacks that conceal harmful intent within complex adversarial prompts. Existing defenses primarily rely on input perturbation or harmful-output suppression, but they rarely model where malicious intent resides, resulting in brittle protection and excessive ov...

Luo-Yu Chen, Wei-Qi Wang, Chen-Han Zhang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.