Open access
Aug 2026
Leveraging Large Language Models to Detect and Revise Unsafe Responses in Context-Sensitive Dialogues
This work proposes a pipeline that leverage LLMs as safety detector, editor and evaluator to mitigate undesired behaviour in human-computer dialogues and shows reduction in the unsafe dialogues after revision.
T. Ajayi, M. Arcan, P. Buitelaar
· WOCHAT2026: Workshop on Chat... · 0 citations