DMART-HAR: Dynamic Multimodal Transformer Learning for Cross-Domain Human Activity Recognition
Abstract
Human activity recognition (HAR) in smart environments plays a critical role in applications such as healthcare monitoring, intelligent transportation systems, and ambient assisted living; however, existing approaches are limited by their inability to effectively handle heterogeneous multimodal sensor data, capture long-range temporal dependencies, and generalize across diverse real-world environments under domain shifts. In this work, we present DMART-HAR, a Dynamic Multimodal Activity Recognition Transformer framework that unifies structured multimodal representation learning, transformer-based temporal modeling, cross-modal interaction, and adversarial domain adaptation within a single architecture. Specifically, the developed method introduces a sensor tokenization mechanism to encode heterogeneous IoT data into a unified representation space, followed by a transformer encoder to capture global contextual dependencies, while a cross-modal attention module enables deep interaction among sensor modalities and an adversarial domain adaptation strategy enhances robustness to unseen environments. Extensive experiments on benchmark datasets, including CASAS, PAMAP2, and Opportunity, demonstrate that DMART-HAR consistently outperforms both conventional baselines and recent state-of-the-art methods, achieving accuracy/F1-scores of 94.3%/92.8%, 96.2%/94.7%, and 89.8%/88.1%, respectively, and consistently outperforms the strongest competing approaches under cross-domain evaluation settings. These findings demonstrate the effectiveness of modeling temporal dynamics, multimodal relationships, and domain invariance simultaneously, establishing DMART-HAR as a scalable and robust solution for real-world HAR applications.