Integrated computer vision system for estimating video parameters using transmission protocols and cloud computing
Abstract
Traditional video surveillance systems often rely heavily on manual analysis, which can delay the detection of relevant events. Although deep learning helps reduce this inefficiency, its high computational cost limits its use in existing edge infrastructures. This paper presents an integrated computer vision system that operates under a distributed (edge-to-cloud) architecture, capable of performing object detection and depth estimation using only 2D video streams. Computation was decentralized by streaming video via RTSP to a GPU in the cloud. For visual perception, different variants of the YOLO model were tested on a own dataset of 5,922 images, while the zeroshot model Depth Anything V2 model was used to estimate relative monocular depth, converting it to absolute distance via polynomial regression. The results highlight the YOLOv12s model, which achieved an precision of 89.7% and a mAP@0.5 of 94.1%. The regression for the spatial metric achieved a coefficient of determination (R2) of 0.9888. This research demonstrates that it is possible to obtain a scalable and cost-effective spatial understanding of basic CCTV infrastructure, without the need for expensive sensors.