Jul 2026· 2026 4th International Conference on Sustainable Computing and Smart Systems (ICSCSS)· pp. 873-878· 0 citations· 20 references
Abstract
Face recognition systems applied to smart surveillance settings often experience poor performance when the faces are partially occluded by a mask or other objects. Occlusions eliminate critical facial information, which makes face identification much more difficult for traditional deep learning models. To solve this issue, a hybrid deep learning model utilizing convolutional neural networks and transformer-based attention mechanism is proposed in this study for robust masked and occluded face recognition. The framework uses the ResNet50 backbone for obtaining the discriminative local facial favorable features, and the Vision Transformer module for obtaining long-range context relationships between facial regions. In addition, an Adaptive Occlusion Attention Module is introduced to Zurcrook visible facial areas and neglect the corrupted features to occlusions. Experiments were carried out on the Real-World Masked Face Dataset (RMFD) with 1205 images of 25 identities. The proposed model attained 93.46% training accuracy and Top-1 and Top-5 recognition accuracy were 51.87% and 81.33%, respectively. Additional occlusion experiments resulted in occlusion recognition accuracy of 28.63% and cross-dataset evaluation using MaskedFace-Net resulted in an average feature similarity of 0.8288. The results show that the proposed hybrid architecture enhances the recognition robustness of masked and partially obstructed facial images facing the surveillance situation.
A comparative analysis of existing studies is presented to highlight the evolution of deep learning techniques and their effectiveness in improving recognition accuracy and computational efficiency and emerging research directions are outlined to provide insights for future research.
Patel Bhautika Ronak· International journal of res...· 0 citations
A Self supervised Capsule Network (SS-CapsNet) integrates convolutional feature extraction, capsule based learning, dynamic routing and contrastive self supervised learning into a combined framework for robust face recognition.
K. Minney Prisilla, N. Jayashri· International journal of com...· 0 citations
An occlusion-aware hybrid biometric framework for reliable 3D face recognition that reaches an accuracy of up to 98.7%, even in partial occlusions, and significantly reduces the Equal Error Rate, demonstrating its effectiveness and suitability for real-world biometric authentication applications.
M. L. Gangadhar, A. S. Raju, C. R. Roopashree· Engineering, Technology &...· 0 citations
Facial expression recognition technology is vital for security, verification, and personalization, but it faces challenges due to variations in scale, illumination, occlusion, and facial expressions. This paper presents a hybrid architecture that combines Vision Transformers (ViTs) to capture global context with EfficientNet-B3 for multi-scale feature extraction. Unlike simple concatenation, our approach projects the ViT’s [CLS] token and the EfficientNet’s global pooling features into a shared 512-dimensional space before merging, enabling better alignment of global and local features. When tested on the FERPlus dataset, it reaches an accuracy of 94.4 ± 0.3%, surpassing several recent methods, notably existing transformer- and CNN-based methods. Ablation studies show each component’s contribution, with the full model outperforming the no-fusion version by 2.6%. With around 98 million parameters and an inference time of ~23 ms per image, it balances efficiency and high performance, suitable for real-time use on suitable hardware. Evaluation via confusion matrix, t-SNE visualization, and comparisons with recent techniques such as HLA-ViT (90.13%), AU-ViT (90.15%), and CCFER (91.24%) demonstrates its robustness and discriminative feature learning. This work highlights the promise of hybrid deep learning architectures in tackling real-world facial expression recognition challenges.
Sasan Karamizadeh, Saman Shojae Chaeikar, Mazdak Zamani· Journal of Imaging· 0 citations
The proposed Adaptive Super-Resolution Generative Adversarial Network (Adaptive SRGAN) integrates adaptive learning with image super resolution to reconstruct identity preserving high resolution facial images by employing adaptive learning rate optimization, dynamic loss weighting, attention guided feature enhancement and identity preserving loss functions.
M. Kirubakaran, A. S. Aneeshkumar· International journal of com...· 0 citations
Facial identity identification in unrestricted real-world environments may benefit from this model, which performs well in identifying and verifying low-quality and cross-pose masked faces and outperforming the various state-of-the-art methods and previously proposed methods.
P. Kaur, Taqdir Kaur, Sahezpreet Singh· Engineering Research Express· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.