Unlocking the head: unleashing deep learning and depth camera for free head movement gaze estimation
Abstract
Gaze estimation has applications such as visual attention analysis and human-computer interaction. However, performance under free head movement conditions can be further enhanced. This study aims to improve the accuracy and robustness of appearance-based gaze estimation without head fixation. We develop a gaze-tracking system using a consumer depth camera and a deep-learning-based gaze estimation model [Gaze-Point-Net (GPN)]. A customized YOLOv5-based detector is developed to simultaneously localize the eyes and mouth, providing reliable facial landmarks for gaze estimation. The detected regions and depth image were used to calculate a vector of location and posture. In GPN, a double-channel convolutional neural network with a Squeeze-and-Excitation module extracts features from binocular images, and these features are concatenated with the location-and-posture vector to complete the multimodal gaze-position prediction task. Our GPN presents competitive performance in four experiments: (1) The error distribution in the sequential point test ranges from 4 to 10 degrees; (2) In the random point test, the calibrated and filtered GPN achieved a pixel error of 185.19 ± 116.57 pixels and an angular error of 4.52° ± 2.87°, demonstrating superior performance over counterpart methods; (3) In the trajectory tracking test, approximately 81.91% of gaze points were located within a 200-pixel tolerance radius of the target trajectory; (4) In the browsing test, the generated gaze heat and trajectory maps showed satisfactory results. The developed eye-tracking system, integrating a depth camera and deep learning models, demonstrates competitive performance and strong potential for several eye-tracking applications.