This paper provides a comprehensive survey of the current state of action recognition, focusing specifically on three open-world challenges: the integration of multimodalities, the ethical and social implications of these technologies, and the utilization of feedback mechanisms to enhance model performance.
Abstract
Action recognition has emerged as a critical area of research within the realm of computer vision, driven by the increasing demand for intelligent human–machine systems capable of understanding and interpreting human behaviors in the real world. The ability to decipher intricate details of human actions holds immense potential to improve system design, predictive modeling, data-informed decision-making, and real-time operational improvements across a wide variety of domains. Some examples of applications range from surveillance and real-time management of public spaces and infrastructure systems, to development of predictive modeling and robotic systems for individualized healthcare interventions, to implementing effective human–computer interaction in both professional and recreational settings. This paper provides a comprehensive survey of the current state of action recognition, focusing specifically on three open-world challenges: the integration of multimodalities, the ethical and social implications of these technologies, and the utilization of feedback mechanisms to enhance model performance. We delve into the evolution of action recognition, from early feature-based approaches to the deep learning revolution, emphasizing how the incorporation of multiple sensory modalities—such as visual, audio, and depth data as well as other cues—has advanced the field. Furthermore, we examine the ethical challenges associated with deploying these technologies in the public domain, particularly regarding privacy, bias, and societal impact, and discuss the need for responsible development and regulation. The third focus of the paper is the use of top-down and bottom-up feedback mechanisms within deep learning architectures, exploring how these strategies can mimic human cognitive processes to improve accuracy and reliability in action recognition systems. By identifying current gaps and proposing future research directions, this paper aims to inspire continued innovation in this dynamic and impactful field for intelligent systems.
This work aims to establish a unified perspective on vision-based mistake analysis in procedural activities, highlighting its potential across diverse domains and aspects and categorizing approaches based on their use of procedural structure, supervision levels and learning strategies.
Konstantinos Bacharidis, Antonis A. Argyros, Hazel Doughty· 0 citations
The integration of artificial intelligence and robotics into clinical medicine is no longer a question of whether but of how, and physicians currently caring for older adults with complex multimorbidity need practical guidance for the transformation ahead. This commentary offers a framework organized around three fundamental domains of clinical care-information collection, data analysis, and treatment delivery-and describes how the physician's role within each is shifting rather than disappearing. In information collection, the clinician moves from direct performer to supervisor, deciding which data streams and alerts warrant attention and eliciting the contextual, values-based history no algorithm can capture. In data analysis, where AI will prove most transformative, the physician becomes a critical appraiser and ethical arbiter of machine-generated options, a role demanding vigilance against documented hazards: training-data bias, digital ageism, large language model confabulation, and automation bias. In treatment delivery, the physician becomes an orchestrator, matching the level of intervention to the patient's goals across a spectrum from robotic nursing support to autonomous self-management. Throughout, the organizing principle is augmented intelligence-AI as an amplifier of expert physician judgment rather than a replacement for it-operationalized as an architectural safeguard in which AI outputs remain interrogable, sourced, and subordinate to the responsible clinician. A specialty-endorsed certification of clinical AI tools, modeled on the American Geriatrics Society Beers Criteria, is proposed. Geriatric medicine's emphasis on multimorbidity, goals-of-care conversations, and interdisciplinary coordination, exemplified by the PACE model, positions geriatricians to lead this transition and to prepare, beginning today.
Richard G. Stefanacci· Journal of The American Geri...· 0 citations
Large Action Models (LAMs) extend the capabilities of AI systems beyond text generation toward perception, reasoning, and action, enabling applications across robotics, autonomous systems, smart manufacturing, healthcare, and the Internet of Things. This Special Issue of ACM Transactions on Multimedia Computing, Communications, and Applications (TOMM) brings together five contributions addressing key challenges in LAM research, including safety and robustness against jailbreak and adversarial attacks, semantic-perceptual integration for robotic manipulation, efficient deployment on edge devices, and natural-language-driven decision-making for networked systems. Together, these papers span the theoretical, implementation, and application dimensions of LAMs, offering both practical solutions and a foundation for future research toward LAM-based systems that are safe, efficient, and reliably grounded in action.
M. Gabbouj, Jin Li, Xin Lin et al.· ACM Transactions on Multimed...· 0 citations
This paper presents a conceptual design for a proactive human assisting robot system capable of recognizing human activities and responding proactively. The system leverages contextual human activity recognition to interpret human actions across diverse contexts, while behavior trees are utilized to define dynamic and interpretable robot behaviors. We outline the system architecture, incorporating contextual human action recognition (HAR), behavior trees (BTs), and ROS, using the Spot robot platform as a representative example. We explain how HAR enables the robot to provide proactive assistance, discuss its limitations, and introduce methodologies for contextual HAR to address these limitations, thereby enhancing the robot's decision-making in complex human activity scenarios.
The rapid development of autonomous technologies such as artificial intelligence (AI), machine learning, and robotics has transformed the relationship between humans and machines. While these systems enhance efficiency and productivity, they raise critical concerns about preserving human values in automated decision-making. Key values such as fairness, accountability, transparency, privacy, and dignity are often overlooked, as autonomous systems rely heavily on data-driven algorithms that may embed biases. This misalignment can lead to discrimination, reduced trust, and unintended consequences in sensitive sectors like healthcare, transportation, and criminal justice. Addressing this issue requires integrating human values into system design through an interdisciplinary approach combining computer science, ethics, sociology, and policy studies. The paper highlights challenges such as value ambiguity, contextual differences, and technical limitations in encoding ethics. It proposes a mixed-method approach involving stakeholder analysis, value elicitation, algorithm auditing, and ethical performance evaluation. Findings suggest that systems incorporating multi-stakeholder input and adaptive ethical constraints align better with human values than purely data-driven models. The study emphasizes the importance of regulatory frameworks, transparency, explainable AI, and continuous monitoring to ensure ethical deployment. Ultimately, embedding human values in autonomous systems is a socio-ethical necessity, requiring scalable, context-aware, and culturally sensitive solutions to achieve sustainable technological development.
Shalini Gupta· International Journal of Inn...· 0 citations
Human activity recognition (HAR) is a well-established research domain in which traditional deep learning approaches often struggle with complex human–object interactions and generalization capabilities. While multimodal large language models (MLLMs) provide strong perceptual encoding for vision-based HAR, they remain limited in handling long-horizon temporal reasoning and maintaining consistency across complex activity sequences. Reinforcement learning (RL) plays a complementary and indispensable role by enabling reward-driven optimization over extended temporal contexts, allowing MLLMs to reason over long-duration activities and adapt to dynamic HAR scenarios. This paper presents a general framework for incorporating pretrained MLLMs into HAR tasks, with a particular focus on the role of RL methods. We systematically review recent advances in the integration of MLLMs and RL for HAR, with a focus on pipeline design, policy optimization methods, reward verification, and training efficiency. Additionally, we summarize applications and existing visual datasets commonly used in HAR research. Finally, we identify key challenges and suggest future research directions to facilitate more generalizable and robust HAR systems empowered by MLLMs and RL.