Jul 2026
Hierarchical Latent Reasoning for LLM-based Recommendation
To further optimize the reasoning trajectory, HiLaR combines final recommendation feedback with layer-aware process rewards derived from the marginal target-likelihood gain of each state, and generally outperforms strong sequential, generative, and LLM-based recommendation baselines.
Peiyu Hu, Siying Gu, Weihai Lu et al.
· arXiv.org · 1 citation