Teaching Machines to See: A Narrative Review of Computer Vision from Roberts's Blocks to Convolutional Depth and Detection
Abstract
Computer vision---making machines interpret images---traveled from blocks-world edge finders to deep convolutional networks matching human benchmarks, and its history is AI's most complete case of representation learning's triumph. This article presents a narrative review of the field's canonical line: Roberts's 1963 machine perception of solids, Marr's 1982 computational vision, Viola and Jones's 2001 face detection, Lowe's 2004 SIFT features, Dalal and Triggs's 2005 HOG descriptors, Felzenszwalb and colleagues' 2010 deformable part models, Szeliski's 2010 synthesis, Girshick's 2015 Fast R-CNN, Long, Shelhamer, and Darrell's 2015 fully convolutional nets, Simonyan and Zisserman's 2015 VGG, He and colleagues' 2016 ResNet, and Redmon and colleagues' 2016 YOLO. The synthesis is organized around three themes: representation, in which hand-engineered features gave way to learned hierarchies; architecture, in which convolution, regions, and residual depth solved recognition's geometry; and tasks, in which classification widened into detection, segmentation, and real-time video. It is concluded that vision's deep learning settlement reorganized the field around data and compute---and that its open problems, robustness and embodiment, define the current frontier.