Jul 2026
How Does Alignment Tuning Shape Representations of Sycophancy and Related Cue-Induced Biases in LLMs?
Cue-induced bias is best understood not as a single flaw in LLMs but as a family of causally effective linear directions that are largely shaped by alignment tuning.
Prakhar Gupta, Terry Jingchen Zhang, Florent Draye et al.
· arXiv.org · 1 citation