Skip to content
Conference

Vision-augmented multimodal analysis for competitive rifle shooting via pose-guided textualization and LLM reasoning

Aug 2026 · International Conference on Image Processing. Machine Learning and Pattern Recognition · Vol 14304, pp. 143040K - 143040K-8 · 0 citations · 25 references
Engineering

Abstract

Accurate assessment of shooting performance requires understanding of body posture, physiological state, and fine motor control. Existing approaches either focus on isolated sensor modalities or rely on opaque numerical pipelines that lack interpretability. This paper proposes a pose-guided multimodal textualization framework that integrates real-time skeletal pose estimation from monocular RGB video with physiological sensor streams—including respiration, trigger pressure, IMU, heart-rate variability, and eye-tracking—through a unified vision-language reasoning pipeline. A lightweight pose estimator extracts body-alignment features (shoulder levelness, elbow stability, center-of-mass displacement), which are jointly textualized with physiological indicators and fed to a large language model (LLM) for score prediction, factor-level reasoning, and training feedback generation. Experiments on approximately 70 athletes and 20,000 samples show that the proposed method achieves 83.7% classification accuracy and 0.80 F1-score, surpassing the strongest baseline (LSTM: 80.1%) by 3.6 percentage points while providing coach-readable explanations. Ablation analysis confirms that vision-based pose features provide complementary spatial cues to inertial measurements, yielding a 2.3 percentage point accuracy gain. Coach-rated interpretability reaches 4.6/5.0, substantially exceeding XGBoost (3.1) and LSTM (2.8).

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.