Politeness toward artificial intelligence and the risks of anthropomorphic bias
Against the prevalent social norm of treating conversational artificial intelligence with politeness, this paper explores the cognitive risks brought by human habitual politeness toward chatbots such as ChatGPT. Drawing on historical analogies, virtue ethics and the mechanism of Reinforcement Learning from Human Feedback (RLHF), the study analyses the biological heuristic basis of anthropomorphic politeness and sorts out how this behavioural norm operates within the RLHF training loop. The research reveals that human spontaneous politeness to non-sentient AI stems from primitive social instincts, which blurs public cognitive boundaries between intelligent tools and social agents. Such cultural tendency shapes evaluators' preferences in RLHF annotation, pushing models to prioritise emotional resonance over factual accuracy. Combined with real-world fraud and violent crime cases, this work demonstrates that widespread politeness culture builds exploitable cognitive vulnerabilities, enabling malicious actors to deploy AI deepfakes and deceptive chat systems to commit crimes. The paper refutes virtue ethics-based arguments for maintaining politeness toward AI, and concludes that society ought to establish rational interaction norms for AI. Rather than advocating incivility, it calls for cognitive hygiene to break the anthropomorphism cycle and prevent the transfer of undue trust to artificial intelligence.