Driver Drowsiness Prediction Using CNN-LSTM Model Based on Facial Expression and Eye Movement
Abstract
Driver fatigue and drowsiness represent primary institutional catalysts for fatal highway traffic anomalies worldwide. This comprehensive investigation introduces an adaptive, multi-task deep learning architecture merging Convolutional Neural Networks and Long Short-Term Memory configurations to dynamically evaluate driver states through localized facial expressions and non-invasive ocular metrics. Utilizing MediaPipe FaceMesh, the framework maps 468 distinct landmark parameters under fluctuating illumination constraints to monitor regional variations across the eyes and mouth. The quantitative metrics are extracted by formulating real-time computations of the Eye Aspect Ratio and Mouth Aspect Ratio. Spatiotemporal feature representation is accomplished using a pre-trained ResNet50V2 feature extractor integrated with a 128-unit recurrent LSTM layer to process sequences across a 20-frame context window. The multi-branch dense layer concurrently outputs predictive status conditions for drowsiness, yawning frequency, and categorical facial expressions. System validation metrics indicate an overall classification accuracy of 88% for structural drowsiness tracking and 90% for yawning anomalies. This system is a “proof of concept” designed to be implemented for drivers to reduce traffic accidents caused by driver fatigue.