A Review of Reinforcement Learning With Multimodal Large Language Models for Vision-Based Human Activity Recognition
Human activity recognition (HAR) is a well-established research domain in which traditional deep learning approaches often struggle with complex human–object interactions and generalization capabilities. While multimodal large language models (MLLMs) provide strong perceptual encoding for vision-based HAR, they remain limited in handling long-horizon temporal reasoning and maintaining consistency across complex activity sequences. Reinforcement learning (RL) plays a complementary and indispensable role by enabling reward-driven optimization over extended temporal contexts, allowing MLLMs to reason over long-duration activities and adapt to dynamic HAR scenarios. This paper presents a general framework for incorporating pretrained MLLMs into HAR tasks, with a particular focus on the role of RL methods. We systematically review recent advances in the integration of MLLMs and RL for HAR, with a focus on pipeline design, policy optimization methods, reward verification, and training efficiency. Additionally, we summarize applications and existing visual datasets commonly used in HAR research. Finally, we identify key challenges and suggest future research directions to facilitate more generalizable and robust HAR systems empowered by MLLMs and RL.