Skip to content

Author

Runyang You

We have 1 of 6 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Jul 2026

SLPO: Scaling Latent Reasoning via a Surrogate Policy

Surrogate Latent Policy Optimization (SLPO) is introduced to bring outcome-reward RL to autoregressive latent reasoners: an empirical surrogate policy density over latent transitions for trajectory-level credit assignment, and a correctness-supervised stopping head that outcome-reward optimization refines into a variable-horizon policy.

Runyang You, Zhiyuan Liu, Yongqi Li et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.