Skip to content
Conference

ActionLMM: captioning long-video actions with memory-augmented VLMs

Aug 2026 · International Conference on Digital Image Processing · Vol 14351, pp. 1435123 - 1435123-9 · 0 citations · 35 references
Engineering

TL;DR

This work introduces ActionLMM, a memory-augmented vision-language model for long-video action summarization that aligns visual and motion modalities through joint representation learning and leverages a novel dual-memory mechanism to retain both local motion details and global temporal structure.

Abstract

The success of large language models (LLMs) has inspired the development of foundation-level multimodal systems that integrate vision and language. However, current video-language models—such as Video-LLaMA and VideoChat—struggle with fine-grained human motion understanding and fail to summarize long videos effectively. Meanwhile, motion-focused models are limited to short clips and lack mechanisms to capture long-range spatiotemporal context. We introduce ActionLMM, a memory-augmented vision-language model for long-video action summarization. It aligns visual and motion modalities through joint representation learning and leverages a novel dual-memory mechanism to retain both local motion details and global temporal structure. To support evaluation, we propose a large-scale benchmark dataset with 33,887 longform action videos and 169,435 caption annotations across 1920 action categories. Experiments show that ActionLMM significantly outperforms prior methods, offering a robust and scalable solution for fine-grained human action understanding.

View source

Similar papers

Jul 2026

TinyMem: Condensing Multimodal Memory for Long-Form Video Action Detection.

This paper introduces TinyMem, a model built upon compact multimodal memory for long-form video action detection that outperforms a range of state-of-the-art models on AVA v2.2 while using 5 times fewer memory tokens than the baseline with dense visual memory embeddings.

Rui Tian, Qi Dai, Hang-Rui Hu et al. · 0 citations
Conference Jul 2026

AVT-PAC: A Pipeline for Multimodal Action Prediction and Captioning

Action prediction from frames and videos is a well-studied problem. Models trained with a single modality, mostly vision, will fail in low-light conditions. Recent works have attempted to predict action categories using vision-language and audio-visual models. A challenge, however, is that some dataset annotations lack temporal ground truth and include only the vision modality. Relying on transformers and an intelligent Vision Language Model (VLM) is a viable solution, but deploying them on edge devices could lead to reduced performance and hallucinations. This work presents an Audio-Visual-Text (AVT)-based multi-step Pipeline for Action Prediction and Captioning (AVT-PAC) to address this problem. First, for an input video, we identify the area to focus on using the Region-of-Interest (ROI) Extraction module. CLIP and CLAP encoders are used for ROI prediction. However, the ROI extracted region may vary in duration, resulting in a large number of frames to be processed. To avoid learning from redundant frames, we uniformly sample key frames within the ROI extracted region using a keyframe extraction module. These keyframes are then used to train an Audio-Visual Action and Text-Aware Representation (AVATAR) model to predict actions and captions. Through systematic experiments, we demonstrated that the proposed AVATAR-TCN model beats the present state-of-the-art (SOTA) baselines on the AVE dataset. Code is available in https://github.com/Ifovia/AVT-PAC

A. R, Ambarish Parthasarathy, Sucharitha Devarakonda et al. · 0 citations
Preprint Aug 2026

UniMem: Unifying Multimodal Memory and Control for Vision-Language-Action Models

UniMem is presented, a framework that unifies high-level, multimodal memory and low-level control under one backbone that outperforms fixed-interval image sampling baselines in simulation and hierarchical baselines in hardware, while offering faster inference and a simple training pipeline for easy adoption.

Lars W. Osterberg, M. Wang, Mac Schwager · 0 citations
Preprint Aug 2026

COMET: Contrastive Motion-Enhanced Temporal Reasoning for Video Multimodal Large Language Models

COMET is a temporally grounded framework that systematically strengthens video MLLMs through explicit temporal representation, appearance-motion fusion, and direction-aware optimization and achieves consistent overall improvements with a pronounced motion-temporal bias.

Chenghua Zhu, Zhaolu Kang, Qifan Shi et al. · 0 citations
Open access 2026

Cross-Modal Temporal Alignment for Action Grounding in Videos

A cross-modal temporal alignment framework that combines a multi-scale temporal convolutional encoder with capsule-based dynamic routing, jointly optimizing temporal boundary prediction, cross-modal semantic alignment, and capsule diversity is proposed.

Gengtian Shi, Chenhao Wu, Shaofei Wang et al. · 0 citations

Condensed Test-Time Adaptation of VLMs for Action Recognition

A novel training-free Condensed Dynamic Adapter C ON DA is proposed, which leverages vision-text alignment to guide vision-vision alignment and is compatible with arbitrary VLM and generalizes well across complex scenarios, such as long-term and egocentric scenarios.

Wenxuan Ge, Hongyu Qu, Rui Yan et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.