Skip to content

Similar papers

Preprint Aug 2026

Measuring and Detecting Harmful AI Sycophancy

It is demonstrated that detection performance drops on unseen models and an initial approach is proposed to address this challenge, and it is shown that detecting PSRS is feasible from the response text alone, and detectors need to learn subtle PSRS patterns from the training data.

Bohan Jiang, Dawei Li, Yasin N. Silva et al. · 0 citations
Review Open access Sep 2026

The Impact of Explanation Characteristics on the Effects of Recommender Systems Explanations

There seems to be a consensus that adding explanations to AI systems, such as recommender systems, can have positive effects, such as increasing a user’s trust, the transparency of the system, or the efficiency of a decision-making process. It is to this date unclear, though, how an explanation needs to be designed to...

Kathrin Wardatzky, Oana Inel, Luca Rossetto et al. · 0 citations
Open access Sep 2026

The Readability of Generative AI Policies: An Empirical Analysis of OpenAI’s Informational Documents

In recent years, media attention has focused on artificial intelligence, particularly on chatbot services and generative intelligence. ChatGPT, created by OpenAI, was one of the earliest online tools and rapidly gained popularity. Users are indeed exposed to a service with privacy notifications and conditions of use th...

Jacopo Bassetta, D. Perpetuini, Maria Teresa Giusti et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Measuring LLM Sycophancy under Sustained Multi-Turn Pressure

By analyzing models with accessible reasoning traces, it is found that the correct position often remains represented in a reasoning trace when the response concedes, suggesting that the model chooses to please a user and sycophancy is not due to lack of knowledge or ignorance.

Lei Tang, Kangda Wei, Tian-Yu Jiang et al. · 1 citation
Preprint Aug 2026

From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop

The Workshop on Trustworthy Natural Language Processing (TrustNLP), co-located with major ACL conferences since 2021, has grown from 8 proceedings papers to 41 over six editions, documenting a field-wide transition from post-hoc interpretability of static models to mechanistic understanding and proactive control of gen...

Rahul Gupta, Abhinav Mohanty, Anaelia Ovalle et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.