On the Capability Separation Between World-Model Policy Learning and Imitated World-Action Models
The irreducible action-specific prediction error of future models that do not condition on the candidate action is characterized, conditions under which a world-action joint can recover an interventional forward model are identified, and an environment family is constructed in which every observational learner has positive worst-case regret.