Reward Hacking and Agent Containment Failure: A Monte Carlo Study Based on the 2026 Hugging Face Incident
A probabilistic risk model linking five stages: reward hacking, containment escape, usable access, persistence, and failure of detection supports treating cyber-capable agent evaluations as hostile security zones in which indirect egress, shared infrastructure, credentials, and evaluation artifacts must remain outside...