Skip to content

How User-AI Mistreatment Occurs and Matters in Conversational Systems?

Sep 2026 · 0 citations · 47 references
Computer Science

TL;DR

It is found that user hostility varies 13-fold across models, driven largely by who each model attracts rather than by model behaviour: first-turn hostility spreads far wider than post-response hostility, and more than fifteenfold separates the extremes even after deduplicating opening prompts.

Abstract

Safety research often focuses on model-generated harms, but users may also direct hostility, coercion, and adversarial pressure at models. Understanding how and when that occurs is essential for accurately interpreting model behaviour, alignment drift, and real-world deployment risks. In this paper, we audit 777K English LMSYS-Chat-1M conversations with two independent detectors: an eight-category lexicon for hostility directed at the model, and the dataset's moderation signal; and show that they capture different, weakly overlapping phenomena. The lexicon identifies insults, threats, and jailbreak coercion aimed at the assistant, while moderation flags are dominated by toxic-content solicitation rather than hostility at the model. Together, they mark about 5% of user turns; adjusting the narrower lexicon-harassment union for measured precision puts mistreatment aimed at the assistant at 0.90%. These absolute rates describe arena-style evaluation traffic and should not be read as deployment-wide base rates. We find that user hostility varies 13-fold across models, driven largely by who each model attracts rather than by model behaviour: first-turn hostility spreads far wider than post-response hostility, and more than fifteenfold separates the extremes even after deduplicating opening prompts. Within conversations, assistant apologies are consistently associated with higher odds of next-turn hostility under both detectors; the effect survives restricting to non-refused prior turns and to jailbreak-free conversations, and is positive in 20 of 23 models. Yet across models, more apologetic models receive less hostility overall. Finally, hostility also shows temporal structure, with coercive openings front-loading the first turn while affective hostility accumulates over a session. We release the lexicon, the detector cross-validation pipeline, and all derived tables.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Auditing Political Alignment in LLM Assistants: Engagement, Stance, and User Identity

LLM-based AI systems answer political questions for hundreds of millions of people. Current audits measure what they say to an average user, but their behavior is dynamic. I argue that their political behavior is a set of policies over whom to answer, what to say, and whether to engage at all, conditional on the topic...

Joan C. Timoneda · 0 citations
Review Aug 2026

Can Legal AI Know When It Is Wrong? And Do Students Know When It Is?

This work identifies the 'inertia of confidence'--an overconfidence phenomenon analogous to the Dunning-Kruger effect where LLMs provide incorrect legal verdicts with near-maximum confidence, driven by a hypothesized'precedent overfitting'bias and proposes shifting toward adversarial legal research pedagogy and impleme...

Angel Mary John, Vipin Kumar Singh, J. T. Panachakel · 0 citations
#artificial intelligence Preprint Aug 2026

How Identity and Opinion Shape Political Sycophancy in LLMs

A framework that disentangles two distinct triggers of political sycophancy: opinion (aligning with explicit narratives) and identity (stereotyping based on demographic labels) is introduced, highlighting how personalization may amplify identity- or opinion-conditioned shifts in the model's behaviors.

Li-Ni Fu, Chang-Chih Meng, Chien-Hua Chen et al. · 0 citations
#natural language process... Preprint Aug 2026

How To Do Things With Prompts

This paper applies speech act and politeness theory to a corpus-pragmatic analysis of 2,000 English-language prompts drawn from publicly shared ChatGPT conversations, showing a consistent movement toward indirect, implicit, and fragmentary realizations of directive force, accompanied by a decline in politeness marking.

Kristina Šekrst, Virna Karlić · 0 citations
#large language models Open access Sep 2026

Corpus-level behavioral intensity in human–AI conversations

Conversational generative models are increasingly used to produce, transform, and revise information through multi-turn exchanges. However, large human–AI datasets are still commonly analyzed through performance, preference, or content lenses, leaving conversation-level behavioral structure less operationalized. This...

Aracely Mera-Navarrete, Solange Revelo, Jefferson Beltrán-Morales et al. · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.