Skip to content
Review Open access

Benchmark Research on Safety Value Alignment Evaluation of Open-Domain Dialogue Systems based on NLP

Aug 2026 · Scientific Journal of Intelligent Systems Research · 0 citations · 21 references

Abstract

Large language models sometimes behave in puzzling ways. They pass various safety tests, yet in multi-turn dialogues, a few carefully crafted sentences can lead them astray into making dangerous judgments. We call this "extreme value alignment failure." A review of recent research reveals an awkward situation: attack, defense, evaluation, and theoretical studies operate in isolation, with little connection among them. This fragmentation results in repeated extreme risks that remain unresolved. This paper maps the four research directions onto a unified framework---"failure mode, attack vector, defense level, evaluation benchmark"---providing a theoretical coordinate for the field and directions for future evaluation research.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.