Self-Confirming Superposition Traps in Reinforcement Learning
It is shown that this loop can sustain a lower-return policy even when representation fitting is globally optimal on data selected by the agent, which then uses the resulting returns to guide its next choices.