A Simple Baseline for Learning Approximate State Abstractions in Factored State Spaces
This work introduces a surprisingly simple neural network architecture change: a learnable, state-independent attention mask applied to the inputs of the policy and value networks and trained end-to-end using only the RL objective.
Anshuman Senapati, Josiah P. Hanna
· 0 citations