Skip to content

Author

Jia-Cheng Shi

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Aug 2026

DocPO: Advancing Document Policy Optimization via Tailored Step-Aware Rewards

This work proposes Step-Aware Annealing (SAA), a plug-and-play reward sharpening mechanism that progressively increases reward curvature during training, amplifying subtle quality differences among high-scoring samples while preserving stability in early learning.

Yunhao Wang, Bing-Hong Wu, Zhenyu Huang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.