It is found that user hostility varies 13-fold across models, driven largely by who each model attracts rather than by model behaviour: first-turn hostility spreads far wider than post-response hostility, and more than fifteenfold separates the extremes even after deduplicating opening prompts.
Abstract
Safety research often focuses on model-generated harms, but users may also direct hostility, coercion, and adversarial pressure at models. Understanding how and when that occurs is essential for accurately interpreting model behaviour, alignment drift, and real-world deployment risks. In this paper, we audit 777K English LMSYS-Chat-1M conversations with two independent detectors: an eight-category lexicon for hostility directed at the model, and the dataset's moderation signal; and show that they capture different, weakly overlapping phenomena. The lexicon identifies insults, threats, and jailbreak coercion aimed at the assistant, while moderation flags are dominated by toxic-content solicitation rather than hostility at the model. Together, they mark about 5% of user turns; adjusting the narrower lexicon-harassment union for measured precision puts mistreatment aimed at the assistant at 0.90%. These absolute rates describe arena-style evaluation traffic and should not be read as deployment-wide base rates. We find that user hostility varies 13-fold across models, driven largely by who each model attracts rather than by model behaviour: first-turn hostility spreads far wider than post-response hostility, and more than fifteenfold separates the extremes even after deduplicating opening prompts. Within conversations, assistant apologies are consistently associated with higher odds of next-turn hostility under both detectors; the effect survives restricting to non-refused prior turns and to jailbreak-free conversations, and is positive in 20 of 23 models. Yet across models, more apologetic models receive less hostility overall. Finally, hostility also shows temporal structure, with coercive openings front-loading the first turn while affective hostility accumulates over a session. We release the lexicon, the detector cross-validation pipeline, and all derived tables.
LLM-based AI systems answer political questions for hundreds of millions of people. Current audits measure what they say to an average user, but their behavior is dynamic. I argue that their political behavior is a set of policies over whom to answer, what to say, and whether to engage at all, conditional on the topic...
This work identifies the 'inertia of confidence'--an overconfidence phenomenon analogous to the Dunning-Kruger effect where LLMs provide incorrect legal verdicts with near-maximum confidence, driven by a hypothesized'precedent overfitting'bias and proposes shifting toward adversarial legal research pedagogy and impleme...
Angel Mary John, Vipin Kumar Singh, J. T. Panachakel· 0 citations
A framework that disentangles two distinct triggers of political sycophancy: opinion (aligning with explicit narratives) and identity (stereotyping based on demographic labels) is introduced, highlighting how personalization may amplify identity- or opinion-conditioned shifts in the model's behaviors.
Li-Ni Fu, Chang-Chih Meng, Chien-Hua Chen et al.· 0 citations
It is found that students'detection accuracy improves over time, driven by a shift from relying on linguistic cues to leveraging shared social and contextual signals.
Dan Schumacher, Pragathi Durga Rajarajan, Haven Kotara et al.· 0 citations
This paper applies speech act and politeness theory to a corpus-pragmatic analysis of 2,000 English-language prompts drawn from publicly shared ChatGPT conversations, showing a consistent movement toward indirect, implicit, and fragmentary realizations of directive force, accompanied by a decline in politeness marking.
Conversational generative models are increasingly used to produce, transform, and revise information through multi-turn exchanges. However, large human–AI datasets are still commonly analyzed through performance, preference, or content lenses, leaving conversation-level behavioral structure less operationalized. This...
Aracely Mera-Navarrete, Solange Revelo, Jefferson Beltrán-Morales et al.· Frontiers in Artificial Inte...· 0 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduSep 2, 2026
A new method, called CW-Net, translates the reasoning process of an autonomous vehicle’s AI system into understandable concepts that explain its behavior.