Predicting Consumer In-Store Purchase Through Real-Time Video Analytics: An Advanced Computer Vision and Deep Learning Approach
Abstract
Physical retailers have long lacked the real-time behavioral visibility that online platforms enjoy through clickstream data. This research addresses that gap by introducing a video analytics framework that transforms in-store security camera footage into a rich, structured behavioral record: an "offline clickstream." Using computer vision and deep learning techniques, including person re-identification, trajectory reconstruction, pose estimation, and vision-language models, the system extracts moment-by-moment signals of shopper intent: how customers move through the store, how they interact with products, and how their body language evolves during a visit. A transformer-based prediction model trained on these signals achieves dramatically better purchase prediction accuracy than conventional demographic or contextual benchmarks alone: improving predictive performance by up to 79% on key metrics. Beyond prediction, the framework supports five real-time targeting policies; simulations show that a persuadability-based policy yields a 13.1% profit lift over no targeting. For retailers and policymakers, this research offers a scalable, privacy-conscious blueprint for bridging the capability gap between physical and digital commerce, enabling timely, personalized interventions that improve customer experience and store profitability.