Skip to content
Open access

ForensicNet: Lightweight Attention-Enhanced MobileNetV2 for Automated Face Identification

N. J. Savitha B. T. Lata
Jul 2026 · Engineering, Technology & Applied Science Research · Vol abs/2607.16273, pp. 37620-37625 · 0 citations · 23 references
Computer Science

TL;DR

The suggested model combines the MobileNetV2 backbone with Convolutional Block Attention Modules (CBAM) to improve the learning of discriminative features while maintaining computational speed, and requires only 2.1 GFLOPs per inference, and hence can be used in real-time forensic surveillance applications.

Abstract

In forensic environments, automated identification of perpetrators is difficult due to pose changes, changes in light, occlusion, and lack of labeled data. This paper presents ForensicNet, a lightweight deep learning framework for forensic face recognition that enhances attention. The suggested model combines the MobileNetV2 backbone with Convolutional Block Attention Modules (CBAM) to improve the learning of discriminative features while maintaining computational speed. A two-phase transfer learning strategy with adaptive layer unfreezing is used to improve domain adaptation and reduce overfitting. This study used publicly available datasets such as LFW and SCFace, with 15,000 facial images spanning 68 identity classes. The proposed model outperforms baseline architectures such as AlexNet, ResNet-50, and MobileNetV2, with an accuracy of 92.4%, a precision of 90.8%, and a recall of 89.5%. Additionally, the framework requires only 2.1 GFLOPs per inference, and hence can be used in real-time forensic surveillance applications.

Read PDF

Similar papers

Open access Aug 2026

Lightweight Face Anti-Spoofing with MobileNetV2: Transfer Learning, Model Compression, and Evaluation on LCC-FASD

Face recognition systems remain vulnerable to presentation attacks involving printed photographs, replayed videos, and three-dimensional masks, creating a need for accurate and computationally efficient face anti-spoofing methods. This study proposes a lightweight deep learning framework based on MobileNetV2 for detecting genuine and spoofed facial presentations. The model employs transfer learning and fine-tuning and is trained and evaluated on the Large Crowd-Collected Face Anti-Spoofing Dataset (LCC-FASD), which contains genuine and spoofed facial samples captured under varying conditions. The framework integrates image preprocessing, data augmentation, and hyperparameter optimisation to improve classification performance and generalisation. Experimental results show a precision of 95.86%, a recall of 98.45%, and an F1-score of 97.88%. Training and validation analyses, confusion-matrix evaluation, and receiver operating characteristic analysis further support the stability of the classification results. Pruning and quantisation are also applied to reduce computational complexity while preserving competitive detection performance. The resulting lightweight framework is suitable for biometric authentication systems, mobile devices, online identity-verification platforms, and edge-based security applications. The study demonstrates the potential of MobileNetV2 for efficient face anti-spoofing and provides a basis for future research on advanced presentation attacks, including deepfakes and three-dimensional masks.   Spanish-language metadata / Metadatos en españolTítulo en español:Detección ligera de ataques de suplantación facial con MobileNetV2: aprendizaje por transferencia, compresión del modelo y evaluación en LCC-FASD Resumen:Los sistemas de reconocimiento facial continúan siendo vulnerables a ataques de presentación mediante fotografías impresas, videos reproducidos y máscaras tridimensionales, lo que genera la necesidad de métodos precisos y computacionalmente eficientes para detectar la suplantación facial. Este estudio propone un marco ligero de aprendizaje profundo basado en MobileNetV2 para distinguir entre presentaciones faciales genuinas y fraudulentas. El modelo emplea aprendizaje por transferencia y ajuste fino, y se entrena y evalúa con el conjunto de datos Large Crowd-Collected Face Anti-Spoofing Dataset (LCC-FASD), que contiene muestras faciales genuinas y fraudulentas capturadas en condiciones diversas. El marco integra preprocesamiento de imágenes, aumento de datos y optimización de hiperparámetros para mejorar el desempeño de clasificación y la capacidad de generalización. Los resultados experimentales muestran una precisión del 95.86 %, una exhaustividad del 98.45 % y una puntuación F1 del 97.88 %. Los análisis de entrenamiento y validación, la evaluación mediante matriz de confusión y el análisis de la característica operativa del receptor respaldan adicionalmente la estabilidad de los resultados de clasificación. Asimismo, se aplican técnicas de poda y cuantización para reducir la complejidad computacional, al tiempo que se mantiene un desempeño de detección competitivo. El marco ligero resultante es adecuado para sistemas de autenticación biométrica, dispositivos móviles, plataformas de verificación de identidad en línea y aplicaciones de seguridad basadas en el borde. El estudio demuestra el potencial de MobileNetV2 para la detección eficiente de suplantación facial y proporciona una base para futuras investigaciones sobre ataques de presentación avanzados, incluidos los deepfakes y las máscaras tridimensionales. Palabras Claves:detección de suplantación facial; detección de ataques de presentación; MobileNetV2; aprendizaje por transferencia; red neuronal convolucional; autenticación biométrica; LCC-FASD. Smart citations: https://scite.ai/reports/10.61467/2007.1558.2026.v17i4.1529Dimensions.Open Alex.

M. Abbas, Muhammad Munib, Muhammad Sajid et al. · 0 citations
Open access Aug 2026

Real-Time Face Mask Detection Using Transfer Learning with MobileNetV2

The COVID-19 pandemic created an urgent need for automated systems capable of verifying face-mask compliance in public spaces. This paper presents a lightweight, real-time face-mask classifier built on transfer learning with the MobileNetV2 architecture. A pretrained ImageNet backbone is used as a frozen feature extractor with a compact classification head, followed by a fine-tuning phase that unfreezes the final convolutional layers. The model is trained and evaluated on the publicly available. Face Mask ∼12K Images dataset, comprising approximately twelve thousand pre-cropped and class-balanced face images split into training, validation, and test partitions. Using data augmentation, two-phase training, and standard regularization, the classifier attains approximately 99% accuracy on the held-out test set with near-perfect precision and recall for both the masked and unmasked classes. The results confirm that a low-compute, mobile-oriented backbone combined with transfer learning is sufficient for accurate binary mask detection, making the approach suitable for deployment on edge devices. The proposed pipeline is a clean, reproducible, end-to-end implementation rather than a novel methodology.

M. Younus · 0 citations
Open access Aug 2026

DEEPFAKE FACE DETECTION IN VIDEOS USING OPENCV AND MOBILENETV2

The broad dissemination of altered facial photographs, especially Deepfakes, which are getting harder to identify with traditional techniques, is made possible by the Internet's quick development. While existing methods concentrate on intricate network architectures or geographical domain properties, they sometimes lack resilience against advanced counterfeit techniques. In order to overcome this, we suggest a Deepfake detection framework based on MobileNetV2, which uses effective convolutional feature extraction to accurately classify real and fake facial photos. In order to guarantee consistent input quality and improve the discriminative features for detection, the framework starts with OpenCV-based preprocessing, which includes face detection, alignment, and normalisation. By automatically learning hierarchical spatial features from the pre-processed facial photos, MobileNetV2, a lightweight yet powerful convolutional neural network, replaces the requirement for manually created features.

A. Mohitha, C. B. Jones · 0 citations
Review Aug 2026

Periocular Soft Biometrics: A Survey and Applications to Multimedia Forensics and Disinformation Detection

A survey of demographic attribute estimation from periocular images, covering publicly available datasets, methodological trends from handcrafted descriptors to deep learning architectures, and the state of the art in gender, age, and ethnicity prediction is provided.

F. Alonso-Fernandez, Kevin Hernandez-Diaz, J. Bigun · 0 citations
#artificial intelligence Preprint Sep 2026

Swin Meets EfficientNet: Lightweight Architectures for GAN-Based Face Forensics

Modern generative models, such as GANs, diffusion architectures, and autoregressive systems, now produce facial images that are nearly indistinguishable from authentic photographs. This capability makes detecting forged images increasingly difficult, raising serious concerns about identity theft, fraud, and misinformation campaigns. Our research focuses specifically on GAN-generated synthetic faces, which underpin many face-centric deepfakes, and investigates efficient detection approaches using image analysis alone. Existing detection systems rely heavily on either convolutional neural networks (CNNs) or global vision transformers. While CNNs excel at identifying texture-based local features, they struggle with broader contextual understanding. Traditional Vision Transformer (ViT) models can capture long-range structures effectively, but demand substantial computational resources. Our work explores Swin-Transformer-based architectures across three implementations: a compact Swin Transformer trained from the ground up, ImageNet-1K pre-trained Swin-Tiny and Swin-Small models adapted for binary classification, and a novel hybrid combining EfficientNet-B0's convolutional processing with a Swin Transformer backend. We evaluated all models using the 140K Real and Fake Faces dataset, which includes StyleGAN-generated fake faces alongside authentic images from Flickr and DFDC, with balanced splits for training, validation, and testing. The EfficientNetB0+Swin hybrid achieved 99% accuracy and a 99.44% recall on 5,000 test images, outperforming both pure Swin variants and a previous CNN-only baseline on this dataset. Our results suggest that combining hierarchical CNN features with shifted-window self-attention provides an efficient and computationally lightweight method for detecting GAN-generated synthetic faces.

S. Basu, Ashima Sood, Vijay Kumar et al. · 0 citations
Jul 2026

Physiological Signals as a Forensic Modality for Talking-Face Deepfake Detection

A detection framework is proposed that extracts per-video rPPG wave- forms via RhythmFormer and trains a suite of lightweight classifiers to distinguish real from synthesized physiologi- cal signals and shows that detec- tion difficulty is strongly method-dependent.

Othmane Harraq, Tamer Aldwairi · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.