Skip to content

Author

Jinghao Yang

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Politeness toward artificial intelligence and the risks of anthropomorphic bias

Against the prevalent social norm of treating conversational artificial intelligence with politeness, this paper explores the cognitive risks brought by human habitual politeness toward chatbots such as ChatGPT. Drawing on historical analogies, virtue ethics and the mechanism of Reinforcement Learning from Human Feedback (RLHF), the study analyses the biological heuristic basis of anthropomorphic politeness and sorts out how this behavioural norm operates within the RLHF training loop. The research reveals that human spontaneous politeness to non-sentient AI stems from primitive social instincts, which blurs public cognitive boundaries between intelligent tools and social agents. Such cultural tendency shapes evaluators' preferences in RLHF annotation, pushing models to prioritise emotional resonance over factual accuracy. Combined with real-world fraud and violent crime cases, this work demonstrates that widespread politeness culture builds exploitable cognitive vulnerabilities, enabling malicious actors to deploy AI deepfakes and deceptive chat systems to commit crimes. The paper refutes virtue ethics-based arguments for maintaining politeness toward AI, and concludes that society ought to establish rational interaction norms for AI. Rather than advocating incivility, it calls for cognitive hygiene to break the anthropomorphism cycle and prevent the transfer of undue trust to artificial intelligence.

Jinghao Yang · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.