2025· International Journal of Intelligent Automation & Robotics Engineering· Vol 8, pp. 01-16· 0 citations
TL;DR
Experimental evaluations demonstrate that the proposed framework outperforms CNN-based and hybrid approaches on standard robotic perception benchmarks, achieving over 98% visual perception accuracy with improved scene understanding, localization, obstacle detection, navigation, and computational efficiency.
Abstract
Autonomous robots play a crucial role in industrial manufacturing, healthcare, transportation, logistics, agriculture, disaster response, planetary exploration, and service robotics. Reliable visual perception is essential for enabling robots to recognize objects, understand scenes, localize themselves, and navigate safely in dynamic environments. Although CNN-based vision models have significantly improved perception accuracy, they often struggle to capture long-range dependencies and generalize to complex or unseen environments. Recent advances in Transformer-based vision models address these limitations by employing self-attention mechanisms to learn both local visual features and global contextual relationships. Architectures such as Vision Transformer (ViT), Swin Transformer, DETR, SAM, and Mask2Former have achieved remarkable performance in object detection, semantic segmentation, SLAM, localization, obstacle avoidance, and autonomous navigation. This paper presents a comprehensive review and proposes the Transformer-Based Visual Perception Models for Autonomous Robots (TBVPM-AR) framework. The framework integrates RGB cameras, depth sensors, LiDAR, IMUs, multimodal sensor fusion, transformer-based feature extraction, contextual reasoning, and edge-cloud computing to achieve robust perception in dynamic environments. Mathematical formulations for self-attention, positional encoding, and feature embedding provide the theoretical foundation of the architecture. Experimental evaluations demonstrate that the proposed framework outperforms CNN-based and hybrid approaches on standard robotic perception benchmarks, achieving over 98% visual perception accuracy with improved scene understanding, localization, obstacle detection, navigation, and computational efficiency. The proposed architecture offers a scalable, explainable, and adaptable solution for future Industry 5.0, collaborative robotics, autonomous vehicles, and smart cyber-physical systems.
Autonomous robotic navigation has become a key capability for intelligent robots operating in dynamic environments such as warehouses, hospitals, smart cities, agriculture, and autonomous transportation. While supervised learning methods achieve strong navigation performance, they depend on large labeled datasets that are costly and time-consuming to obtain. Self-Supervised Learning (SSL) addresses this limitation by enabling robots to learn robust visual and spatial representations directly from unlabeled sensor data through self-generated learning objectives. This paper reviews recent advances in SSL techniques, including representation learning, contrastive learning, predictive learning, masked image modeling, and multimodal sensor fusion for autonomous navigation. It also examines the integration of data from RGB cameras, LiDAR, IMUs, GPS, depth sensors, and odometry to improve perception, localization, obstacle avoidance, and path planning in unknown environments. Finally, the paper discusses key challenges such as domain adaptation, computational efficiency, safety, and continual learning, highlighting SSL's potential to enable scalable, adaptive, and lifelong autonomous robotic navigation.
D. Michie, Roger Needham· International Journal of Int...· 0 citations
Robotic object recognition is a fundamental capability that enables autonomous robots to interact intelligently with dynamic environments. Traditional vision-based methods, such as SIFT, SURF, HOG, and template matching, perform well under controlled conditions but struggle with variations in lighting, viewpoint, occlusion, and complex backgrounds. Recent advances in deep learning have significantly improved recognition accuracy by automatically learning features from raw image data. However, individual deep learning models often face challenges related to computational cost, inference speed, and limited generalization. This paper proposes a Hybrid Deep Learning Framework that integrates Convolutional Neural Networks (CNNs), Vision Transformers (ViTs), attention mechanisms, and multimodal sensor fusion (RGB, depth, and LiDAR) to enhance recognition accuracy and efficiency. The framework combines local and global feature extraction, adaptive feature fusion, intelligent object recognition, and robotic decision-making for real-time perception and task execution. It supports applications in industrial automation, warehouse logistics, autonomous mobile robots, healthcare, agriculture, and service robotics while improving robustness, scalability, and computational efficiency.
Seshagiri N, Mahabala H. N.· International Journal of Int...· 0 citations
We propose a novel methodology for autonomous landing zone detection in Micro Aerial Vehicles (MAVs) based on Vision Transformers (ViTs). The core contribution of this work lies in demonstrating that transformer-based architectures, originally developed for large-scale vision tasks, can be effectively adapted to safety-critical aerial robotics applications with limited training data. Unlike traditional Convolutional Neural Networks (CNNs), ViTs leverage self-attention mechanisms to model long-range spatial dependencies, enabling a more holistic understanding of scene geometry and surface suitability for landing. We systematically evaluate the proposed approach on aerial RGB images from a public dataset as well as on noisy depth images captured onboard a drone using a lightweight depth camera. Our results show that the ViT-based model consistently outperforms widely used CNN architectures, including ResNet, particularly in low-data regimes where generalization is crucial. Notably, the transformer model maintains strong robustness even when operating on degraded depth inputs. In addition to accuracy improvements, the proposed system achieves real-time performance, with an average inference time of [Formula: see text] ms on legacy GPU hardware. These findings highlight the practical feasibility and effectiveness of Vision Transformers for reliable, efficient MAV landing zone detection.
Victoria Eugenia Vazquez-Meza, J. Martínez-Carranza· Unmanned Systems· 0 citations
An intelligent vision-based autonomous robotic framework that integrates deep learning-based object detection with hybrid adaptive navigation for dynamic environments is proposed in this research. The proposed system addresses the challenges of real-time perception and robust navigation in unstructured settings by combining a convolutional neural network (CNN) for object detection with a hybrid control mechanism for motion planning. The CNN, implemented using a state-of-the-art architecture such as YOLO, processes visual input to identify obstacles and target objects, providing critical environmental awareness. Moreover, the hybrid navigation strategy merges reactive obstacle avoidance, achieved through algorithms like the Vector Field Histogram (VFH), with adaptive path planning using Rapidly-exploring Random Trees (RRT) to ensure both immediate collision avoidance and long-term goal convergence. The integration of these components enables the robotic system to dynamically adjust its navigation policy in response to environmental changes, thereby improving robustness and adaptability. The novelty of our approach lies in the seamless fusion of vision-based perception and adaptive control, which enhances the system’s capability to operate in complex, dynamic scenarios. Experimental validation demonstrates the effectiveness of the framework in real-world applications, highlighting its potential for deployment in autonomous vehicles, service robotics, and industrial automation. The proposed method offers a scalable and efficient solution for autonomous systems requiring high levels of situational awareness and adaptive decision-making.
K. A.· International Journal on Rob...· 0 citations
A systematic analysis of the machine learning and deep learning models underpinning vehicle autonomy, spanning classical convolutional neural networks for object detection and semantic segmentation to recurrent and Transformer-based architectures for trajectory prediction and motion planning is presented.
Esraa Khatab, Fares Fathy, Abdallah AlKholy et al.· Machine Learning and Knowled...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.