Preprint
Jul 2026
Steering Robustness into World Action Models via Mechanistic Interpretability and Optimal Control
This work uses mechanistic interpretability to study how robustness-relevant perturbations are represented in WAM activation space and predicts strong steerability in the Cosmos-Policy and DiT4DiT models but weak steerability in LingBot-VA, consistent with steering intervention results.
Jihoon Hong, Julian Skifstad, Qiyu Dai et al.
· 1 citation