Basic Characteristics of the Four Visual Perception Models of UAV
Pure visual perception of Unmanned Aerial Vehicle (UAV) is the core supporting technology to realize autonomous flight. In recent years, with the deep integration of visual perception and artificial intelligence methods, related technologies are rapidly evolving from the target detection level to the environmental semantic understanding level. Focusing on the intelligent development of pure visual perception models, this paper systematically summarizes the four mainstream technology routes of convolutional neural network (CNN), transformer, CNN‑transformer hybrid architecture, and visual‑language model, focusing on the integration mechanism of AI key technologies such as attention mechanisms and multimodal alignment. The research shows that lightweight CNN is still the engineering foundation of UAV airborne real-time sensing. CNN‑transformer hybrid architecture significantly improves the adaptability to complex scenes. The visual‑language model makes it easier to realize the paradigm transition of open vocabulary recognition and semantic navigation. At present, the core challenges in the field focus on the prominent contradiction between model intelligence and airborne edge computational force constraints, and the lack of robustness in a dynamic environment. In the future, lightweight multimodal architecture and end-to-end‑edge‑cloud collaborative computing will become the key development direction for UAVs to achieve fully autonomous intelligent flight.