Large language models frequently fail to balance staying truthful with being supportive. They often exhibit sycophancy in responses to users, agreeing with false claims, offering unwarranted flattery, and giving advice skewed toward users'expressed views. In reality, sycophancy rarely happens in a single exchange; it m...
Sidharth Pulipaka, R. Binkytė, Ivaxi Sheth et al.· 0 citations
This work introduces artifact-mediated propagation, where adversarial content introduced through an artifact is stored in an assistant's persistent memory, reproduced in a subsequently created artifact, and acquired by another assistant that later reads it.
Sidharth Pulipaka, Anshu Sharma, Stanislau Hlebik et al.· 0 citations
Large language models (LLMs) are increasingly deployed with external tools that extend what they can do beyond their own knowledge. Tools help on tasks that need external information, but their availability may also change how a model handles questions that do not need them. Prior work has mostly asked whether models s...
Saanvi Paturi, Arsen Kenzhebayev, Arham Sethi et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.