Skip to content

Author

Seonvin Cho

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#machine learning Preprint Sep 2026

Multi-step Proximal Policy Improvement in Offline Reinforcement Learning

Offline reinforcement learning (RL) must reconcile two competing requirements: policy updates should stay near dataset-supported actions to keep value estimates reliable, yet meaningful gains often require moving beyond the behavior distribution. We develop a geometric view of offline actor updates by modeling policies as a probability manifold endowed with a chosen metric geometry. Under this lens, a broad class of offline actor objectives can be interpreted as a single proximal policy improvement step (SPI), i.e., an implicit discretization of a manifold gradient flow induced by a critic-defined energy. Building on this insight, we propose multi-step proximal policy improvement (MPI), a plug-in refinement mechanism that composes sequential re-centered proximal steps. MPI enables controlled policy improvement beyond dataset support while retaining proximal control at each refinement. The framework accommodates multiple policy geometries and admits practical instantiations for deterministic and diagonal-Gaussian policies. Experiments on D4RL benchmarks show that small numbers of MPI refinements improve strong offline baselines, including TD3+BC, ReBRAC, and IQL, on many tasks. Focused diagnostics further distinguish re-centered refinement from fixed-objective update scheduling and characterize limitations under critic error.

Soohyun Choi, Seonvin Cho, Songnam Hong · 0 citations
#machine learning Preprint Aug 2026

PathBridger: Subgoal Bridges for Offline Goal-Conditioned Reinforcement Learning

The proposed PathBridger is a hierarchical offline GCRL method that explicitly connects subgoal selection to short-horizon execution, and constructs a state-space bridge toward the selected intermediate endpoint and decodes it into a short executable action chunk using an inverse dynamics model.

Soohyun Choi, Seonvin Cho, Songnam Hong · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.