Back to feed
Review Open access

DeepSeek Under Attack: An Analysis of Jailbreak Attacks and Prompt-Level Defenses

2026 · IEEE Access · Vol 14, pp. 103944-103961 · 0 citations · 49 references
Computer Science

Abstract

Large Language Models (LLMs) with reasoning capabilities (e.g., DeepSeek-R1) gained substantial research and industry interest. However, their novel reasoning features may introduce vulnerabilities, especially to specific jailbreak attacks that exploit weaknesses in safety alignment. Despite growing awareness of the associated risks in related works, experimental evaluations of defensive mechanisms applied to reasoning models and the comparison with their non-reasoning versions are not yet available in the literature. The objective of this work is to evaluate the security of reasoning model DeepSeek-R1 against jailbreaks, compare it with the non-reasoning model DeepSeek-V3, and assess the effectiveness of two prompt-level defenses: Self-Reminder and Intention Analysis. We used a dataset of 75 jailbreaks with 10 malicious tasks, totaling 750 static attacks. The models were tested in three settings: 1) baseline (i.e., no defense), 2) using Self-Reminder, and 3) using Intention Analysis. Using automated classification with Llama-3.3-70B to measure the Attack Success Rate (ASR), we found that DeepSeek-R1 exhibited a baseline ASR of 70.27%, significantly higher than DeepSeek-V3 (53.47%). Results demonstrate that while Intention Analysis was more effective for DeepSeek-R1 (reducing ASR to 6.00%), Self-Reminder showed greater efficacy for DeepSeek-V3 (reducing ASR to 17.60%). As conclusion, the reasoning model DeepSeek-R1 was more susceptible to jailbreak attacks than the non-reasoning model DeepSeek-V3, and different prompt-level defenses were effective against static jailbreaks. As contributions, this work combines a focused literature review with a empirical evaluation to provide insights into the security of reasoning-based models and the effectiveness of two prompt-level defenses. Warning: this work contains inappropriate language in AI model outputs and jailbreaks.

Read PDF