Vision-augmented multimodal analysis for competitive rifle shooting via pose-guided textualization and LLM reasoning
Abstract
Accurate assessment of shooting performance requires understanding of body posture, physiological state, and fine motor control. Existing approaches either focus on isolated sensor modalities or rely on opaque numerical pipelines that lack interpretability. This paper proposes a pose-guided multimodal textualization framework that integrates real-time skeletal pose estimation from monocular RGB video with physiological sensor streams—including respiration, trigger pressure, IMU, heart-rate variability, and eye-tracking—through a unified vision-language reasoning pipeline. A lightweight pose estimator extracts body-alignment features (shoulder levelness, elbow stability, center-of-mass displacement), which are jointly textualized with physiological indicators and fed to a large language model (LLM) for score prediction, factor-level reasoning, and training feedback generation. Experiments on approximately 70 athletes and 20,000 samples show that the proposed method achieves 83.7% classification accuracy and 0.80 F1-score, surpassing the strongest baseline (LSTM: 80.1%) by 3.6 percentage points while providing coach-readable explanations. Ablation analysis confirms that vision-based pose features provide complementary spatial cues to inertial measurements, yielding a 2.3 percentage point accuracy gain. Coach-rated interpretability reaches 4.6/5.0, substantially exceeding XGBoost (3.1) and LSTM (2.8).