Skip to content

Semantic Isomorphism Attacks and Defense Evaluation for Jailbreaking Large Language Models

2026 · IEEE Transactions on Audio, Speech, and Language Processing · Vol 34, pp. 4303-4318 · 0 citations · 43 references

Abstract

The safety alignment of large language models (LLMs) faces persistent challenges from jailbreak attacks. While existing methods mostly leverage prompt engineering or adversarial optimization, we identify and formalize an underexplored semantic isomorphism vulnerability where harmful and safe scenarios share highly consistent underlying operational principles, allowing harmful intent to bypass semantic refusal. We present Safe2Harm, a four-stage semantic isomorphism jailbreak attack framework integrating safe rewriting, topic mapping, safe response generation, and topic inversion. Evaluations on 8 mainstream LLMs and three benchmarks show Safe2Harm substantially outperforms state-of-the-art baselines, reaching 99.5% peak ASR on large-parameter models. Further mechanistic and defense analysis reveals the strong input-side stealthiness of the attack, alongside detectable signals in inversion and final outputs, offering key insights for multi-layered LLM safety defense.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.