Skip to content
Open access

Geometry and mask aware vision transformer for masked face recognition in unconstrained scenarios

Aug 2026 · Engineering Research Express · Vol 8 · 0 citations · 52 references
Physics

TL;DR

Facial identity identification in unrestricted real-world environments may benefit from this model, which performs well in identifying and verifying low-quality and cross-pose masked faces and outperforming the various state-of-the-art methods and previously proposed methods.

Abstract

The concealed facial features make it difficult to identify masked faces. The features are further distorted and deteriorated when masked faces are combined with low resolution and pose variation in unrestricted contexts. Pose variation and low-resolution circumstances, along with face masks, are not assessed for current masked face recognition systems. We developed a transformer-based model a mask-aware geometry-guided vision transformer (MGViT), to address these issues. First, a learnable geometry-guided patch weighting (LGW) is used in the proposed model to suppress occluded regions and concentrate on the key face regions. Second, a mask-aware feature adaptor is created to improve the domain embeddings between masked and unmasked faces. Following that, by combining Identity and Consistency Loss functions to align identity, a strong consistency learning is integrated. Experiments with various challenges are carried out on the various masked face datasets. The proposed model, MGViT, performs well in identifying and verifying low-quality and cross-pose masked faces. Additionally, the model achieves 94.70%, 95.25%, 87.67%, and 78.56% accuracy on RMFRD, Masked LFW, Masked Multi-PIE, and Masked LR, respectively, outperforming the various state-of-the-art methods and previously proposed methods. Facial identity identification in unrestricted real-world environments may benefit from this model.

Read PDF

Similar papers

Review Open access 2026

Masked Face Restoration Using GANs: A Survey on Recognition, Detection, and Inpainting

Emergence of masked face recognition (MFR) as a pivotal area in biometric identification has been significantly accelerated by the global COVID-19 pandemic. In response, the research community has developed a variety of innovative techniques to address recognition and detection under occlusion, with a growing emphasis on Generative Adversarial Networks (GANs) for masked face restoration and inpainting. We examined three interconnected sub-domains: Masked Face Recognition (MFR), Face Mask Detection, and Face Unmasking (FU), each addressing unique aspects of the problem from identifying individuals with partially or fully covered faces to reconstructing occluded facial regions for improved accuracy. The core focus of this paper is on the role of GANs in overcoming occlusion by synthesizing realistic facial textures in the masked regions, thereby restoring the identity cues. Beyond technical developments, the paper analyzes the limitations and open research problems, such as maintaining identity consistency in restored images, handling diverse mask types and occlusion levels, and ensuring generalizability across different demographic groups and environments. By integrating insights from recent advances and identifying existing research gaps, this survey aims to serve as a comprehensive reference for academics and practitioners engaged in the development of robust, privacy-aware, and ethically responsible masked face recognition systems enhanced by GANs.

Payal Parekh, Hina Choksi, Mahesh Goyani et al. · 0 citations
Conference Jul 2026

A Hybrid CNN–Transformer Network for Robust Masked and Occluded Face Recognition in Smart Surveillance Systems

Face recognition systems applied to smart surveillance settings often experience poor performance when the faces are partially occluded by a mask or other objects. Occlusions eliminate critical facial information, which makes face identification much more difficult for traditional deep learning models. To solve this issue, a hybrid deep learning model utilizing convolutional neural networks and transformer-based attention mechanism is proposed in this study for robust masked and occluded face recognition. The framework uses the ResNet50 backbone for obtaining the discriminative local facial favorable features, and the Vision Transformer module for obtaining long-range context relationships between facial regions. In addition, an Adaptive Occlusion Attention Module is introduced to Zurcrook visible facial areas and neglect the corrupted features to occlusions. Experiments were carried out on the Real-World Masked Face Dataset (RMFD) with 1205 images of 25 identities. The proposed model attained 93.46% training accuracy and Top-1 and Top-5 recognition accuracy were 51.87% and 81.33%, respectively. Additional occlusion experiments resulted in occlusion recognition accuracy of 28.63% and cross-dataset evaluation using MaskedFace-Net resulted in an average feature similarity of 0.8288. The results show that the proposed hybrid architecture enhances the recognition robustness of masked and partially obstructed facial images facing the surveillance situation.

R. R, Anbalagan E · 0 citations
Open access Aug 2026

A Reliability-Based Multimodal Framework for 3D Face Recognition under Occlusion

An occlusion-aware hybrid biometric framework for reliable 3D face recognition that reaches an accuracy of up to 98.7%, even in partial occlusions, and significantly reduces the Equal Error Rate, demonstrating its effectiveness and suitability for real-world biometric authentication applications.

M. L. Gangadhar, A. S. Raju, C. R. Roopashree · 0 citations
2026

DiEL: Disentangled Evolutionary Learning for Identity-Preserving Face Enhancement and Recognition

The purpose of face enhancement tasks is to improve the recognition of faces, thus adapting to diverse visualization and recognition demands. However, the performance of the majority methods is drastically degraded under extreme conditions, including large pose variations, low resolution, blur, occlusion, and illumination changes, which can distort facial geometry and identity related details. In this work, we construct a simple and effective face robust enhancement method. In particular, in order to maintain the identity consistency of the reconstructed face, an evolutionary learning framework for face disentanglement representation is proposed, in which we disentangle the identity and pose information of the face and unite it with identity recognition as a multi-objective optimization problem, where reconstruction, adversarial, and identity-preserving objectives are adaptively balanced. Further, in order to maintain the pose consistency of reconstructed faces, we construct a unified face pose dictionary, which forms a robust and standard pose representation by statistics and induction of the geometric structure of a large number of face images. In the conditional generation architecture, the pose dictionary could accurately guide the model to realize face reconstruction with desired poses. Extensive benchmark experiments on MS1M, LFW, CPLFW, CFP-FF, CFP-FP, and AgeDB show that the proposed method not only significantly outperforms state-of-the-art methods, but also can further stimulate the discrimination potential of existing face recognition models. Specifically, DiEL achieves an average improvement of 4.66% over the SOTA methods across six benchmark datasets, with particularly significant gains on challenging cross-pose benchmarks such as CPLFW and CFP-FP.

Jingwei Xin, Tian Yang, Jun Hao et al. · 0 citations
Open access Sep 2026

Pyramid-guided multi-scale self-attention and channel–spatial refinement for occlusion-robust face recognition

Facial occlusion degrades face recognition by creating scale-inconsistent identity cues across visible regions and amplifying responses to irrelevant occluders. To address these coupled problems, this study proposes a pyramid-guided multi-scale attention framework based on scale alignment and reliability-aware feature refinement. Hierarchical features are first projected into a shared semantic space, after which local, intermediate, and global dependencies are adaptively weighted according to the available facial information. Channel–spatial refinement is then used to suppress unreliable responses from occluded regions, while an angular-margin objective preserves inter-identity separability from incomplete facial evidence. Experiments on CASIA-WebFace and occluded LFW show that the proposed method achieves an accuracy of 99.26% under clean conditions and 82.72% under occlusion, with an ROC-AUC of 0.8351. At 80% occlusion, the proposed method outperformed the strongest recent baseline, HMPA-GFAF, by 1.41% points and the direct Inception-ResNet-v1 + ArcFace baseline by 5.88% point. These results demonstrate that the proposed framework improves occlusion robustness without sacrificing clean-face recognition performance, indicating its practical potential for identity verification and access-control applications involving masks, glasses, and other partial facial occlusions.

Unknown authors · 0 citations
Conference Jul 2026

Vision Transformer with Attention Rollout for Deepfake Face Image Detection and Localization

Generative AI and synthetic media generation tools have enabled widespread media manipulation tools and raised important privacy concerns with misinformation, identity fraud and the verification of authenticity of media. Most of the current convolution-based deepfake detection methods are hard to be deployed in real scenarios and hard to be interpretable, especially because they have limited ability to capture long-range spatial dependency. It introduces an explainable deepfake face image detection framework based on a vision transformer network for performing powerful binary classification of manipulated and real facial images and an explainable face image localization framework for localizing deepfake image faces. The proposed system involves a transformer-encoder backbone for extracting features through a patches-wise process, which proves suitable for modeling the subtle changes of features when the processes of manipulating the image are designed. To make the network more interpretable, and aid the understanding of the transformer attention distribution as well as localization of manipulated facial regions, a dedicated attention rollout mechanism is embedded. A dedicated rollout mechanism for attention distribution of the transformer and heatmap generating and attention spatial localization are incorporated to improve the interpretability of the network. The framework comprises an end-to-end inference pipeline, such as image preprocessing, estimation of confidence scores, fake-real classification, generation of explainable visualization and storage of prediction history using an integrated database system. An experimental evaluation shows the system can effectively detect deepfakes while also providing accurate local information as justification for classification decisions, contributing to transparency, reliability and trust towards automated synthetic media detection systems.

K. Phani, Shaik Mahaboob, Jailan PG Student et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.