This work proposes Targeted Anti-obfuscation with Mechanistic Enforcement (TAME), which uses Sparse Autoencoders (SAEs) to combine behavioral feedback with targeted suppression of template-associated activations during RL, providing a path from behavioral monitoring to representation-level oversight for more auditable...
Xu-Tao Mao, Jianing Zhu, Jin-Man Zhao et al.· 0 citations
SAGE (Stability-Aware Gradient Efficiency), which maintains difficulty-stratified candidate pools refreshed during training and selects pairs within each pool by a forward-pass signal-to-curvature score, outperforms full-data and size-matched baselines while producing substantially smoother optimization trajectories.
Hui Wu, Hengyi Cai, Jin-Man Zhao et al.· 0 citations
SPACE induces two-level programmatic skills from successful trajectories, where subskill boundaries serve as direct chunk-boundary supervision from trajectory-induced programmatic skills, and distilled into a primitive-chunk policy via hybrid on-/off-policy optimization.
Yan-Ting Yang, Can Jin, Jin-Man Zhao et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.