Jul 2026· International journal of computer information systems and industrial management applications· Vol 18, pp. 861-873· 0 citations
TL;DR
Comparative analysis with state-of-the-art methods including DeepFace, FaceNet, VGGFace, SphereFace, and baseline ArcFace validates the effectiveness of the proposed attention-guided approach for unconstrained face recognition tasks.
Abstract
Face recognition in unconstrained environments remains a challenging problem in computer vision due to variations in pose, illumination, expression, and occlusion. This paper proposes a novel attention-enhanced ArcFace-based deep learning framework that integrates a Residual CNN backbone with Convolutional Block Attention Module (CBAM) and ArcFace loss for robust face recognition. Unlike existing approaches that rely on large-scale external pretraining datasets, the proposed framework is trained exclusively on the Labelled Faces in the Wild (LFW) dataset, demonstrating data-efficient learning. The system is evaluated on both 1:1 verification and 1:N identification protocols. Experimental results demonstrate superior performance with verification accuracy of 95.70%, identification accuracy of 89.75%, ROC-AUC of 99.16%, and True Positive Rate (TPR) of approximately 92% at a 1% False Positive Rate (FPR). The novelty lies in the synergistic integration of attention mechanisms with angular margin-based metric learning, achieving competitive performance without external pretraining. Comparative analysis with state-of-the-art methods including DeepFace, FaceNet, VGGFace, SphereFace, and baseline ArcFace validates the effectiveness of the proposed attention-guided approach for unconstrained face recognition tasks.
A comparative analysis of existing studies is presented to highlight the evolution of deep learning techniques and their effectiveness in improving recognition accuracy and computational efficiency and emerging research directions are outlined to provide insights for future research.
Patel Bhautika Ronak· International journal of res...· 0 citations
Face recognition systems applied to smart surveillance settings often experience poor performance when the faces are partially occluded by a mask or other objects. Occlusions eliminate critical facial information, which makes face identification much more difficult for traditional deep learning models. To solve this issue, a hybrid deep learning model utilizing convolutional neural networks and transformer-based attention mechanism is proposed in this study for robust masked and occluded face recognition. The framework uses the ResNet50 backbone for obtaining the discriminative local facial favorable features, and the Vision Transformer module for obtaining long-range context relationships between facial regions. In addition, an Adaptive Occlusion Attention Module is introduced to Zurcrook visible facial areas and neglect the corrupted features to occlusions. Experiments were carried out on the Real-World Masked Face Dataset (RMFD) with 1205 images of 25 identities. The proposed model attained 93.46% training accuracy and Top-1 and Top-5 recognition accuracy were 51.87% and 81.33%, respectively. Additional occlusion experiments resulted in occlusion recognition accuracy of 28.63% and cross-dataset evaluation using MaskedFace-Net resulted in an average feature similarity of 0.8288. The results show that the proposed hybrid architecture enhances the recognition robustness of masked and partially obstructed facial images facing the surveillance situation.
R. R, Anbalagan E· 2026 4th International Confe...· 0 citations
A Self supervised Capsule Network (SS-CapsNet) integrates convolutional feature extraction, capsule based learning, dynamic routing and contrastive self supervised learning into a combined framework for robust face recognition.
K. Minney Prisilla, N. Jayashri· International journal of com...· 0 citations
Face recognition systems are widely used in surveillance, biometric authentication, access control, and digital identity verification; however, supervision sensitivity, evaluation stability, and performance consistency across datasets remain insufficiently understood. This study investigates the behavior of convolutional, transformer-based, and hybrid face recognition architectures under both Softmax and ArcFace supervision using five-fold subject-disjoint cross-validation on the Labeled Faces in the Wild (LFW) and FAGEv2 datasets. ResNet50, MobileNetV3, DeiT-Small, and a Hybrid multi-branch architecture integrating complementary convolutional and transformer feature representations were evaluated using Top-1 identification accuracy, Area Under the ROC Curve (AUC), Equal Error Rate (EER), computational complexity, and fold-level statistical analysis. Experimental results revealed substantial supervision sensitivity across architectures and datasets. On the LFW dataset, Hybrid-Softmax achieved the highest Top-1 identification accuracy (62.4%), while DeiT-Small-Softmax achieved the strongest verification performance with an AUC of 0.905 and EER of 0.159. On the FAGEv2 dataset, Hybrid-Softmax and DeiT-Small-Softmax achieved the highest identification accuracy (38.0%), while Hybrid-Softmax achieved the strongest verification performance with an AUC of 0.825 and EER of 0.251. Fold-level analyses demonstrated that the effect of ArcFace supervision varied across architectures and datasets, with consistent improvements observed for some convolutional architectures but not for transformer-based or hybrid models. Cross-dataset evaluation further revealed changes in model ranking and supervision behavior, indicating that comparative performance is strongly influenced by dataset characteristics and evaluation conditions. The findings demonstrate that additive angular margin supervision does not universally outperform conventional Softmax optimization and highlight the importance of multi-dataset benchmarking, fold-level evaluation, and supervision sensitivity analysis for robust and reproducible face recognition benchmarking.
Andisani Nemavhola, C. Chibaya, Serestina Viriri· Frontiers in Artificial Inte...· 0 citations
Facial expression recognition technology is vital for security, verification, and personalization, but it faces challenges due to variations in scale, illumination, occlusion, and facial expressions. This paper presents a hybrid architecture that combines Vision Transformers (ViTs) to capture global context with EfficientNet-B3 for multi-scale feature extraction. Unlike simple concatenation, our approach projects the ViT’s [CLS] token and the EfficientNet’s global pooling features into a shared 512-dimensional space before merging, enabling better alignment of global and local features. When tested on the FERPlus dataset, it reaches an accuracy of 94.4 ± 0.3%, surpassing several recent methods, notably existing transformer- and CNN-based methods. Ablation studies show each component’s contribution, with the full model outperforming the no-fusion version by 2.6%. With around 98 million parameters and an inference time of ~23 ms per image, it balances efficiency and high performance, suitable for real-time use on suitable hardware. Evaluation via confusion matrix, t-SNE visualization, and comparisons with recent techniques such as HLA-ViT (90.13%), AU-ViT (90.15%), and CCFER (91.24%) demonstrates its robustness and discriminative feature learning. This work highlights the promise of hybrid deep learning architectures in tackling real-world facial expression recognition challenges.
Sasan Karamizadeh, Saman Shojae Chaeikar, Mazdak Zamani· Journal of Imaging· 0 citations
An occlusion-aware hybrid biometric framework for reliable 3D face recognition that reaches an accuracy of up to 98.7%, even in partial occlusions, and significantly reduces the Equal Error Rate, demonstrating its effectiveness and suitability for real-world biometric authentication applications.
M. L. Gangadhar, A. S. Raju, C. R. Roopashree· Engineering, Technology &...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.