Skip to content
Conference Open access

Incomplete Prompt Jailbreaks in Large Language Models

May 2026 · Annual Meeting of the Association for Computational Linguistics · pp. 31352-31368 · 0 citations · 33 references
Computer Science

TL;DR

This work formalizes this phenomenon as incomplete prompt jailbreaks (IPJ) and provides a systematic empirical characterization of when and how incomplete prompts elicit harmful continuations, and identifies two functional neurons: termination and continuation neurons.

Abstract

Large language models (LLMs) are increasingly released as open-weight models with safeguards against harmful requests. Nevertheless, sentence completion remains vulnerable to incomplete harmful prompts. In this work, we formalize this phenomenon as incomplete prompt jailbreaks (IPJ) and provide a systematic empirical characterization of when and how incomplete prompts elicit harmful continuations. We analyze diverse attractor types associated with incomplete sentence continuation and show that LLMs systematically delay refusal until sentence termination. We further demonstrate that training models to refuse incomplete harmful prompts via parameter tuning is insufficient, failing to generalize across both content domains and attractor types. To enable fine-grained control, we identify two functional neurons: termination and continuation neurons. By clarifying their roles in sentence completion, we highlight the potential of neuron-level interventions for more precise and robust IPJ defenses.

Read PDF

Similar papers

Explaining Jailbreaks: Structured and Interpretable Safety Assessment for Large Language Models

This work proposes an explanation-aware safety framework that augments binary harmfulness detection with structured, human-interpretable explanations capturing severity, strategies, trigger spans, ratio-nales, and derived safety factors, and introduces a human–LLM hybrid annotation and canonicaliza-tion pipeline.

Sunghee Dong, Sungwon Yi, Kangmin Bae et al. · 0 citations
2026

Random Character-Level Perturbations Amplify LLM Jailbreak Attacks

This work finds that models cannot reliably reconstruct the original meaning and layer-wise probe classifiers fail to detect the harmful intent of perturbed prompts, and perturbations can occasionally reduce attack success by inducing off-topic or incoherent responses.

Shuyi Yu, Zhe Cao, Kohei Tsuji et al. · 0 citations
Book Open access Jul 2026

TAPE-JB: Trait-Aligned Prompt Evolution for Jailbreaking Large Language Models

TAPE-JB (Trait-Aligned Prompt Evolution for Jailbreaking LLMs), a genotype-guided evolutionary framework for adversarial prompt generation that searches over structured prompt traits while preserving the original harmful intention is introduced, highlighting the value of structured evolutionary search for systematic red-teaming of language models.

Karolina Seweryn, Anna Wróblewska, Szymon Łukasik · 0 citations
Preprint Aug 2026

Mood Matters: How Syntactic Sensitivity Undermines Safety Alignment

Large language models typically undergo post-training to align them with safety policies but there exist many sophisticated jailbreaks that sidestep established safeguards. For instance, prior work by Andriushchenko et al. (2025) has found that changing the grammatical tense from present to past can be enough to elicit harmful responses. In this work, we uncover a more general failure of non-imperative syntactic forms. We demonstrate that this syntactic vulnerability exists in 16 models up to 70B parameters, using behavioral evaluation. To investigate the root cause, we apply causal mediation analysis, finding that refusal is partially conditioned on upstream syntactic features. By steering these purely syntactic features we are able to trigger and suppress refusal. Finally, we trace this ill-conditioning to linguistically biased post-training data of open-source models and show that increasing syntactic diversity can mitigate the issue. Our findings suggest that current alignment approaches introduce confounders that prevent a pure semantic grounding of the refusal decision.

Alina Klerings, Jannik Brinkmann, Heiner Stuckenschmidt et al. · 0 citations
Preprint Aug 2026

Lexical Perturbations Disrupt LLM Reasoning: An Empirical Study of Attention Diversion

Attention Diversion explains why inference-time strategies, including chain-of-thought prompting, spell-checking, self-repair, and stronger repair models, fail to consistently recover performance: each addresses one channel at a time.

Jiaqi Zhu, Yang Zhang, Junhua Ding et al. · 0 citations
Preprint Jul 2026

Minionese: Comprehensive Benchmark and Mechanistic Study of Multilingual LLM Safety

It is shown that English-only safety evaluations are insufficient; they require accounting for script family, perturbation type, and per-language alignment coverage, and a geometric mechanistic analysis of refusal failure across language tiers.

Chigozirim Ifebi, Brent Kong, Ayushi Mehrotra · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.