Skip to content

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Conference Aug 2026

Lightweight dynamic gesture recognition method for complex background based on improved transformer

In human-computer interaction scenarios, gesture recognition technology enables device operators to perform tasks in a more flexible, natural, and immersive manner. However, most existing gesture recognition algorithms are trending toward lightweight architectures. While this development facilitates real-time performance requirements on mobile devices or resource-constrained systems, it also presents several limitations. Traditional lightweight paradigms (such as simple channel reduction, model pruning, or conventional lightweight convolutions) are often achieved at the expense of global feature representation capabilities, which leads to a sharp decline in recognition accuracy when models face multimodal dynamic hand gesture interaction scenarios with complex backgrounds, occlusions, or non-uniform lighting. Addressing the issue of insufficient global modeling capability, this paper proposes a Vision Transformer architecture based on an efficient additive self-attention mechanism and lightweight dilated convolutions. The method comprehensively absorbs the global modeling capability of the self-attention mechanism and the local feature extraction capability of convolutions, successfully achieving global context modeling capabilities comparable to the standard Transformer utilizing a linear computational complexity cost. To address the aforementioned challenges, this paper proposes a Vision Transformer architecture based on an efficient additive self-attention mechanism and lightweight dilated convolution. The proposed approach integrates the global modeling capabilities of self-attention mechanisms with the local feature extraction capabilities of convolutional networks through optimized combination, achieving a balance between resource utilization and recognition efficiency on mobile platforms. Experimental results demonstrate that the proposed method achieves state-of-the-art accuracy rates of 87.55% on the public dataset NVGestures and 97.92% on the multimodal dynamic gesture dataset Briareo. Furthermore, it achieves comparable or even superior accuracy on the more complex NNGestures dataset through both single-modal and multimodal experiments with reduced model parameters, validating the effectiveness of our methodology.

Huiming Wu, Kun Wang · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.