Skip to content

4 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#artificial intelligence Preprint Sep 2026

Difficulty-Adaptive Tree-Structured Policy Optimization for Expanding Reasoning Coverage in RLVR

Reinforcement Learning with Verifiable Rewards (RLVR) has been central to the recent success of Large Reasoning Models. However, while RLVR significantly improves single-sample accuracy, it often fails to expand the model's intrinsic reasoning coverage (pass@k) due to limited exploration during training. To address this, we optimize the structural design of train-time rollouts to enhance pass@k. Our analysis identifies three key design principles: (1) difficulty-adaptive rollout can play an important role in expanding pass@k, beyond serving as an efficiency heuristic; (2) tree-based rollout outperforms parallel sampling in discovering correct answers; and (3) sentence-entropy-guided forking overcomes the localization phenomenon of token-level branching to maximize semantic diversity. Building on these insights, we propose DATPO (Difficulty-Adaptive Sentence-entropy-guided Tree-structured Policy Optimization). DATPO integrates difficulty-adaptive tree search with a sibling-diversity advantage term, explicitly promoting semantic diversity to expand reasoning coverage during training. Experiments on mathematical reasoning benchmarks demonstrate that DATPO outperforms baselines especially in pass@k, which directly translates to superior test-time scaling performance.

Young Kyu Yu, Sanghwan Jang, Hwanjo Yu · 0 citations
Open access Aug 2026

Personalized federated recommendation via long-horizon local optimization and regularized knowledge guidance

A model-agnostic framework that forms personalized item embeddings through Long-Horizon Local Optimization and injects common global knowledge through intermittent Regularized Knowledge Guidance is proposed and Adaptive Guidance is introduced to control the influence of global knowledge at the user–item interaction level.

Jaehyung Lim, Wonbin Kweon, Woojoo Kim et al. · 0 citations
Preprint Aug 2026

GOD: Enhancing Generalization via Deep Grafting for Sequential Recommendation

Graft-Oriented Distillation (GOD) is proposed, a component-level distillation framework for improved generalization through grafting, which uses selected frozen-teacher components with trainable student counterparts to build hybrid source models.

Woojoo Kim, Junyoung Kim, Jaehyung Lim et al. · 0 citations
#machine learning Preprint Aug 2026

Personalized and Multi-View Representation for Federated Cold-Start Recommendation

PMFRec learns a personalized representation generator to produce user-specific item representations from attribute features, and introduces a global multi-view encoder with item-adaptive gating and an orthogonality objective to capture complementary semantic views while reducing cross-view redundancy.

Jaehyung Lim, Wonbin Kweon, Woojoo Kim et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.