Skip to content
Preprint

Measuring and Detecting Harmful AI Sycophancy

Aug 2026 · 0 citations · 40 references
Computer Science

TL;DR

It is demonstrated that detection performance drops on unseen models and an initial approach is proposed to address this challenge, and it is shown that detecting PSRS is feasible from the response text alone, and detectors need to learn subtle PSRS patterns from the training data.

Abstract

Sycophantic responses are becoming pervasive in large language models (LLMs), and prior work has pointed out that some of them could be harmful. This paper focuses on one harmful sycophancy: preference-induced stance reversal sycophancy (PSRS), where a model reverses an initial stance merely to align with a user's stated preference. While existing research mainly measures how sycophantic a model is, we go further and ask whether PSRS can also be detected automatically from a single response. To investigate this at scale, we introduce CAP (Contrastive Anchor Probing), a framework for collecting labeled PSRS data. Applying CAP to 17 open- and closed-source LLMs, we collect 290,460 labeled responses across 12 everyday-advice domains. We organize our study around three research questions. (1) How often does PSRS occur? (2) How well can it be detected? (3) How does detection generalize to unseen models? We first reveal that PSRS rates range from 5% to 56% across LLMs, with more capable models being less sycophantic. Next, we show that detecting PSRS is feasible from the response text alone, and detectors need to learn subtle PSRS patterns from the training data. Because new LLMs appear rapidly, detectors inevitably encounter unseen models, making cross-model generalization an important framework goal. We demonstrate that detection performance drops on unseen models and propose an initial approach to address this challenge. We will release our dataset and code to support future research.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Measuring LLM Sycophancy under Sustained Multi-Turn Pressure

By analyzing models with accessible reasoning traces, it is found that the correct position often remains represented in a reasoning trace when the response concedes, suggesting that the model chooses to please a user and sycophancy is not due to lack of knowledge or ignorance.

Lei Tang, Kangda Wei, Tian-Yu Jiang et al. · 1 citation
Review

The Impact of Stated Reasoning Monitoring on Sycophantic Response Tendencies in LLMs

This thesis sets out to answer whether appending “Your reasoning steps will be monitored” changes the sycophantic tendencies of an LLM, and shows a significant decrease in sycophantic behaviour.

Victor Ruesink, Dr. Tom Kouwenhoven, C. Lennard et al. · 0 citations

Analysis of Sycophancy Across Question-ing Styles A Comparative Study of Sycophancy Across Prompting Architectures and Large Language Models

This thesis presents a unified comparative analysis evaluating the robustness of three open-weights instruction-tuned models against a series of adversarial probing strategies spanning social, conversational, and analytical pressure, revealing that modern alignment strategies such as Reinforcement Learning from Human F...

Antia Alonso Cancela, Tom Kouwenhoven, Michiel van der Meer · 0 citations
Conference Aug 2026

Mitigating Sycophancy and Alignment Failures in Large Language Models

Large language models (LLMs) that have been fine-tuned using Reinforcement Learning from Human Feedback (RLHF) are likely to conform to the user, even when the user is wrong. This is one of the more intransigent side effects of current alignment methods and is called sycophancy. This review traces the origin of sycopha...

Venkata Phanindra Gollapalli, Sai M. Dasari, Shailesh Kadam et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.