Skip to content
Open access

Self-Supervised Capsule Network for Robust Face Recognition

Aug 2026 · International journal of computer information systems and industrial management applications · 0 citations

TL;DR

A Self supervised Capsule Network (SS-CapsNet) integrates convolutional feature extraction, capsule based learning, dynamic routing and contrastive self supervised learning into a combined framework for robust face recognition.

Abstract

Face recognition becomes one of the most adopted biometric technics due to its applications in intelligent surveillance, access control, border security, digital authentication, criminal investigation and human computer interaction. The development of Deep convolutional neural networks (CNNs) significantly improved accuracy of recognition even in unconstrained environments such as pose variations, illumination changes, facial expressions, occlusions and low-resolution images. Conventional CNNs mainly focus on local spatial features and so it has limited ability to preserve hierarchical association between facial components. Capsule Networks (CapsNets) overcome this by representing visual features as vector capsules with existence and geometric properties of objects. The self supervised learning  of CapsNets enables feature learning from unlabelled images. This article presents a Self supervised Capsule Network (SS-CapsNet) integrates convolutional feature extraction, capsule based learning, dynamic routing and contrastive self supervised learning into a combined framework for robust face recognition. This approach simultaneously learns discriminative identity from large scale unlabelled dataset while preserves facial geometry. The SS-CapsNet provides improved robustness against pose variations, illumination changes, facial occlusions and image degradation.

Read PDF

Similar papers

Open access Jul 2026

Application of Image Enhancement Techniques in Facial Recognition System

The proposed Adaptive Super-Resolution Generative Adversarial Network (Adaptive SRGAN) integrates adaptive learning with image super resolution to reconstruct identity preserving high resolution facial images by employing adaptive learning rate optimization, dynamic loss weighting, attention guided feature enhancement and identity preserving loss functions.

M. Kirubakaran, A. S. Aneeshkumar · 0 citations
Review Open access 2026

Deep Learning-Based Face Detection, Feature Extraction, and Face Recognition from Video: A Comprehensive Review

A comparative analysis of existing studies is presented to highlight the evolution of deep learning techniques and their effectiveness in improving recognition accuracy and computational efficiency and emerging research directions are outlined to provide insights for future research.

Patel Bhautika Ronak · 0 citations
Open access Jul 2026

Attention-Enhanced ArcFace-Based Deep Learning Framework for Unconstrained Face Recognition

Comparative analysis with state-of-the-art methods including DeepFace, FaceNet, VGGFace, SphereFace, and baseline ArcFace validates the effectiveness of the proposed attention-guided approach for unconstrained face recognition tasks.

Samadhan S. Ghodke, Prapti D. Deshmukh · 0 citations
Conference Jul 2026

A Hybrid CNN–Transformer Network for Robust Masked and Occluded Face Recognition in Smart Surveillance Systems

Face recognition systems applied to smart surveillance settings often experience poor performance when the faces are partially occluded by a mask or other objects. Occlusions eliminate critical facial information, which makes face identification much more difficult for traditional deep learning models. To solve this issue, a hybrid deep learning model utilizing convolutional neural networks and transformer-based attention mechanism is proposed in this study for robust masked and occluded face recognition. The framework uses the ResNet50 backbone for obtaining the discriminative local facial favorable features, and the Vision Transformer module for obtaining long-range context relationships between facial regions. In addition, an Adaptive Occlusion Attention Module is introduced to Zurcrook visible facial areas and neglect the corrupted features to occlusions. Experiments were carried out on the Real-World Masked Face Dataset (RMFD) with 1205 images of 25 identities. The proposed model attained 93.46% training accuracy and Top-1 and Top-5 recognition accuracy were 51.87% and 81.33%, respectively. Additional occlusion experiments resulted in occlusion recognition accuracy of 28.63% and cross-dataset evaluation using MaskedFace-Net resulted in an average feature similarity of 0.8288. The results show that the proposed hybrid architecture enhances the recognition robustness of masked and partially obstructed facial images facing the surveillance situation.

R. R, Anbalagan E · 0 citations
Open access Jul 2026

Hybrid CNN with angular margin supervision for robust face identification and verification

Face recognition systems are widely used in surveillance, biometric authentication, access control, and digital identity verification; however, supervision sensitivity, evaluation stability, and performance consistency across datasets remain insufficiently understood. This study investigates the behavior of convolutional, transformer-based, and hybrid face recognition architectures under both Softmax and ArcFace supervision using five-fold subject-disjoint cross-validation on the Labeled Faces in the Wild (LFW) and FAGEv2 datasets. ResNet50, MobileNetV3, DeiT-Small, and a Hybrid multi-branch architecture integrating complementary convolutional and transformer feature representations were evaluated using Top-1 identification accuracy, Area Under the ROC Curve (AUC), Equal Error Rate (EER), computational complexity, and fold-level statistical analysis. Experimental results revealed substantial supervision sensitivity across architectures and datasets. On the LFW dataset, Hybrid-Softmax achieved the highest Top-1 identification accuracy (62.4%), while DeiT-Small-Softmax achieved the strongest verification performance with an AUC of 0.905 and EER of 0.159. On the FAGEv2 dataset, Hybrid-Softmax and DeiT-Small-Softmax achieved the highest identification accuracy (38.0%), while Hybrid-Softmax achieved the strongest verification performance with an AUC of 0.825 and EER of 0.251. Fold-level analyses demonstrated that the effect of ArcFace supervision varied across architectures and datasets, with consistent improvements observed for some convolutional architectures but not for transformer-based or hybrid models. Cross-dataset evaluation further revealed changes in model ranking and supervision behavior, indicating that comparative performance is strongly influenced by dataset characteristics and evaluation conditions. The findings demonstrate that additive angular margin supervision does not universally outperform conventional Softmax optimization and highlight the importance of multi-dataset benchmarking, fold-level evaluation, and supervision sensitivity analysis for robust and reproducible face recognition benchmarking.

Andisani Nemavhola, C. Chibaya, Serestina Viriri · 0 citations
Open access Aug 2026

Enhancing face recognition privacy through the integration of differential privacy and convolutional neural network

The research offers an in-depth analysis of various DP techniques to construct a secure face recognition system employing a Convolutional Neural Network and face classifiers, and concludes that the DP blur with Logistic regression predictors provides the highest privacy, achieving excellent accuracy rates of 97% and 77% for these datasets.

Muhammad Minoar Hossain, Mohammad Motiur Rahman · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.