Polite but Misaligned: Evaluating LLM Politeness Judgments Against Human Pragmatic Norms
Despite strong performance on standard benchmarks, it remains unclear whether large language models (LLMs) evaluate social pragmatics in ways that align with human judgments. We evaluate LLM politeness judgments using two English-language datasets with complementary annotation formats: continuous human ratings and thre...