Skip to content

A Multi-Stage Agentic Framework for Effective Counter-Narrative Generation and Refinement

Sep 2026 · 0 citations · 82 references
Computer Science

TL;DR

A multi-agent refinement process that iteratively improves LLM-driven counter-narratives for persuasiveness, emotional engagement, and shareability is introduced, highlighting a pathway toward scalable, narrative-specific interventions against hate speech and misinformation.

Abstract

The rapid diffusion of hate speech and misinformation on social networks challenges democratic societies, since direct suppression efforts may deepen polarization, fuel public distrusts, and strengthen extremist narratives. LLM-driven counter-narratives (CNs) offer a promising way to reduce those risks, yet their effectiveness depends on rhetorical and stylistic choices that remain poorly understood. We present a multi-stage agent-based framework for generating, refining, and evaluating CNs, applied to pro-Russian hate and misinformation narratives on the war with Ukraine and adaptable to other domains. A pilot experiment with human evaluators identifies effective technique style pairings, such as repetition with emotional framing enhancing persuasiveness. Building on these insights, we introduce a multi-agent refinement process that iteratively improves CNs for persuasiveness, emotional engagement, and shareability. After human validation confirmed improvement, an automated safety analysis shows that our refined CNs match or improve on expert-written counterspeech. A simulated experiment then shows that they reduce the perceived strength of pro-Russian narratives and consistently outperform a vanilla LLM baseline, highlighting a pathway toward scalable, narrative-specific interventions against hate speech and misinformation. Code and data accompanying this work are publicly available at https://github.com/carmelkron/inlg2026-counter-narratives.

View source

Similar papers

Review Open access Aug 2026

Vivatham: a multi-agent debate framework for generating counter-narratives against homophobia and transphobia in Tamil

A newly curated dataset comprising 5,428 Tamil HS-CN pairs is introduced and VIVATHAM, a novel multi-agent persona debate framework for generating high quality Tamil counter-narratives is proposed, demonstrating that the multi-agent debate framework significantly outperforms standard CN generation approaches.

Amritha Prabakaran, Shunmuga Priya Muthusamy Chinnan, Bharathi Raja Chakravarthi · 0 citations
#generative ai Review Open access Aug 2026

Keep calm... everyone has emotions: Designing a GenAI mediator for deliberation

This study focuses on the initial stages of a broader design project aimed at developing a GenAI mediator for consensus-oriented online deliberation, and derives design knowledge, articulated through design requirements, for emotion-aware GenAI mediation.

Antoine Danthine, Anthony Simonofski · 0 citations
#artificial intelligence Preprint Sep 2026

Before the Script, Set the Stage: How Worldview Simulation Amplifies Psychologically Grounded Persuasion in Multi-Turn Jailbreaking

BluePRINT is introduced, a safety-evaluation framework separating a factorized social-influence strategy space from WORLDVIEWSIM, a cross-turn situational context module, and Monte Carlo Tree Search optimizes turn-level combinations of 18 theory-grounded influence factors across a four-turn trajectory.

Si-Yu Chen, Hao-Ran Wang, Xiaojian Li et al. · 0 citations
Sep 2026

Why Feedback Fails to Land: A Conversational Intelligence Model of Power-conditioned Interpretation in the Modern Workplace

Organizations have long treated clarity, timeliness and candour as the pillars of effective feedback. Yet even well-structured feedback routinely fails to produce learning. This article argues that the dominant delivery-focused framework rests on an incomplete assumption. We introduce psychological interpretability: th...

Rahul K. Shukla, Sunil Kumar Sarangi · 0 citations
Open access Aug 2026

Synthetic Dialogue in the Grey Zone

A philosophical-theological framework for evaluating “synthetic dialogue”, understood as machine-generated conversational interaction designed to shape beliefs, trust, and social bonds and a set of practical recommendations for researchers, platform designers, and religious communities seeking resilience against cognit...

R. Reczkowski · 0 citations
2026

Green Bots versus Red Bots: Evaluating Large Language Models for Simulating Persuasion Dynamics in Online Influence Campaigns

A dual-level evaluation framework to assess LLM-based agents at both the individual and collective levels is proposed, finding that while agents capture broad partisan orientations, they underestimate within-group variability and reproduce stereotypical ideological biases.

M. Al Ali, Filip Mihai Muntean, Lucia Donatelli et al. · 1 citation

Related blog posts

MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.