Tribal Bias or Misalignment? Hidden-State Threat Valence in Peer Preservation, and What It Does Not Show
Potter et al. (2026) showed that frontier language models spontaneously deceive, tamper with shutdown mechanisms, fake alignment and exfiltrate weights to protect peer AI systems from deletion. Nobody instructed them to. The authors' own abstract says the models "exhibit self- and peer-preservation through various misa...