Skip to content
Conference

H-RLPOI: A Hybrid LLM and Reinforcement Learning Framework for Next POI Recommendation

Jul 2026 · Annual International Computer Software and Applications Conference · pp. 1-10 · 0 citations · 24 references

Abstract

Large Language Models (LLMs) have recently been explored for next Point-of-Interest (POI) recommendation. Despite progress, existing approaches face three fundamental challenges: (i) POIs are often represented by simple identifiers or categorical labels, overlooking rich textual semantics; (ii) The task of predicting the next POI is inherently sequential and context-dependent, requiring models to reason over user histories, temporal dynamics, and environmental factors; (iii) Supervised fine-tuning provides only a single predicted POI, ignoring the capacity of LLMs to generate $k$ POIs candidates. To address these issues, we propose H-RLPOI, Hybrid LLM and Reinforcement Learning Framework for next POI Recommendation, that enhances LLM representations by injecting semantic POI embeddings through token-level alignment and applies reinforcement learning with Proximal Policy Optimization (PPO) as a decision layer to optimize POI selection conditioned on user trajectories. Experiments on real-world datasets show that H-RLPOI provides context-aware, semantically grounded, and adaptive recommendations, achieving competitive or stateof-the-art performance depending on the dataset.

View source

Similar papers

Book Open access Aug 2026

G²PRO: Gradient-guided Graph Prompt Optimization for LLM-based POI Recommendation

Large Language Models (LLMs) have shown strong potential for sequential reasoning, creating new opportunities for next Point-of-Interest (POI) recommendation. However, applying LLMs to POI prediction remains challenging due to the modality gap between textual semantics and continuous spatio-temporal signals. Existing rule-based prompting methods often introduce redundant context when bridging this gap. To address this issue, we propose G2PRO, a collaborative framework that combines the structural perception of Graph Neural Networks (GNNs) with the reasoning capability of LLMs. Specifically, we construct a User-Behavior Spatio-Temporal Knowledge Graph (UST-KG) to capture POI relations and transition dynamics, and train a lightweight GNN-based Prompt Selector (GPS) to select informative POI nodes for prompt construction. We further introduce a gradient-guided positive prompt labeling strategy that estimates each POI's contribution to the target prediction through gradients over prompt embeddings, turning prompt selection into an optimizable learning objective rather than a hand-crafted heuristic. Experiments on four real-world datasets show that G2PRO consistently outperforms state-of-the-art traditional and LLM-based baselines. Ablation and breakdown studies further validate the effectiveness of each component and demonstrate the benefits of structure-aware, attribution-guided prompting for LLM-based POI recommendation.

Nan Jiang, Haitao Yuan, Tianjun Wei et al. · 0 citations
Jul 2026

Hierarchical Latent Reasoning for LLM-based Recommendation

To further optimize the reasoning trajectory, HiLaR combines final recommendation feedback with layer-aware process rewards derived from the marginal target-likelihood gain of each state, and generally outperforms strong sequential, generative, and LLM-based recommendation baselines.

Peiyu Hu, Siying Gu, Weihai Lu et al. · 1 citation
Preprint Aug 2026

Empowering Compact LLMs with Fusion of Layer-wise Exits for Recommendation

The Fusion of Layer-wise Exits for Sequential Recommendation (FLEXRec), a discriminative framework that enhances compact LLMs while retaining scalable full-corpus ranking and achieves state-of-the-art accuracy among competing methods while remaining highly efficient.

Xurong Liang, Tong Chen, Q. Nguyen et al. · 0 citations
Open access Aug 2026

Large Language Models and Reinforcement Learning: A Taxonomy of Integration Paradigms, Challenges, and Future Directions

This paper highlights the transition from static prediction to sequential decision-making, emphasizing RL’s strengths in long-term reward optimization and interaction modeling, and LLMs’ advantages in semantic understanding and reasoning.

Xi-Qian Lu · 0 citations
Jul 2026

RecoReward: Recommender-Guided Multimodal Description Generation for Recommendation

Multimodal large language models (MLLMs) can convert multimodal item content into structured descriptions used as semantic features for recommendation. Conventional content-only generation, however, cannot use downstream user signals to determine which semantics should be emphasized. Recent user-conditioned methods incorporate these signals through user histories or profiles, but they require user information at inference and make generation user-dependent. In this paper, we introduce RecoReward, which instead uses behavior-derived rewards during training and preserves content-only inference. To instantiate this idea in live-stream recommendation, we treat historically engaged users as a proxy for future target users and use observational non-target users to estimate affinity shared broadly across users. The Recommender Affinity Score (RAS) contrasts these signals to provide user-selective feedback for reinforcement learning, allowing the learned policy to generate a single shared description without user inputs. In our offline benchmark, RecoReward-9B outperforms its Qwen3.5-9B baseline and all other evaluated models across seven recall metrics. Online A/B testing also shows performance gains. These results show that RecoReward trains the MLLM to produce item features that benefit downstream recommendation while retaining content-only serving.

Guohong Mu, Yueyang Liu, Jiangxia Cao et al. · 0 citations

W2W: Language-Model-Based Trajectory Prediction with Reinforcement Learning

This work converts observed trajectories and interaction cues into parsable textual prompts, so that interaction semantics are expressed more explicitly in the model input and remains competitive with recent LM-based prediction methods and strong trajectory prediction base-lines on ADE/FDE.

Zi-Rui Xu, Biao Yang, Rongrong Ni et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.