Are Human-Aligned Models Models of Humans? A Turing-Test Gap in Preference Alignment
It is shown that preference alignment preserves the human response distribution only under a restrictive condition, and no consistent evidence that real human preferences satisfy it, and human-likeness is established as an explicit dimension of alignment rather than something assumed to follow from preference alignment...