It is concluded that effective deepfake governance requires defense in depth integrating forensic detection, verifiable provenance, and institutional accountability.
Abstract
Deepfakes, synthetic audiovisual content produced by deep generative models, have escalated into a critical threat across civilian and military domains, enabling identity fraud, disinformation campaigns, and evidence fabrication. In high-stakes environments, ranging from journalism and finance to healthcare and legal contexts, the consequences extend to severe misinformation, market manipulation, identity fraud, and the erosion of institutional trust. This entry explores how modern visual intelligence and computer-vision techniques are used to detect deepfakes. It outlines key deepfake generation models, such as GANs, autoencoders, neural rendering, and diffusion systems, while also explaining how adversarial methods enhance realism and challenge existing detectors. The overview highlights visual artifacts, digital patterns, and physiological cues commonly leveraged in detection and reviews major CNN, transformer, and frequency-based approaches. It also summarizes evaluation practices and the difficulty of achieving strong generalization. Finally, it identifies emerging directions, including modern intelligence techniques for civilian and military content verification. This survey covers generation architectures (GANs, latent diffusion, neural rendering, video synthesis), the spatial, temporal, frequency-domain, and physiological artifacts they produce, and the detector families that exploit them. We examine evaluation benchmarks and protocols, highlighting cross-generator generalization as the field’s central open challenge. Beyond detection, we discuss cryptographic provenance standards, watermarking, and regulatory frameworks (EU AI Act, DSA, GDPR). We conclude that effective deepfake governance requires defense in depth integrating forensic detection, verifiable provenance, and institutional accountability.
An in-depth survey of fifteen state-of-art methodologies including classical CNN models, temporal-spatial video recognition, transformer-based networks, explainable AI (XAI) models, and models that combine multimodal large language model (LLM) products are provided.
Shavnam Shavnam, Neha Dhiman· International Journal of Inn...· 0 citations
This review presents a comprehensive analysis of recent deep learning and transfer learning techniques for fake image detection, examining widely adopted convolutional neural network architectures, benchmark datasets, evaluation metrics, and current research developments.
Nisha Parveen, Anjali Saxena· International Journal for Re...· 0 citations
Deep-OCR (DeepSeek-OCR) advances document recognition by treating the visual modality as an optical compression medium, enabling long-context OCR at low token cost. However, its increased complexity may introduce new security vulnerabilities. In this paper, we present, to the best of our knowledge, the first pure black-box adversarial attack against a generative OCR vision-language model, where only the decoded string can be queried and no gradients, logits, or model internals are available. We recast the attack as a zeroth-order optimization problem driven by a bounded scalar loss defined directly on the string output via sequence similarity, and estimate the gradient with a random-direction finite-difference scheme whose query cost is independent of the image dimension. An Adam update with ell_infinity projection yields imperceptible perturbations for both untargeted and targeted objectives. Pilot experiments on Deep-OCR validate the string-only attack and evaluation pipeline and expose severe qualitative decoder failures, including repetition, truncation, and prompt leakage. They also show that controlled targeted rewriting remains substantially harder than untargeted degradation; we avoid claiming targeted success until the pre-registered evaluation is complete.
Wenbo Sun, Hong-Zong Li, Yanyun Wang et al.· 0 citations
The rapid advancement of Artificial Intelligence (AI) has facilitated the creation of highly realistic deepfakes and synthetic media, transforming the digital communication landscape while simultaneously generating unprecedented legal and ethical challenges. Deepfakes employ deep learning algorithms, particularly Generative Adversarial Networks (GANs), to manipulate or fabricate images, videos, and audio recordings that closely resemble authentic content. Although these technologies offer legitimate applications in education, entertainment, healthcare, journalism, accessibility, and digital content creation, their misuse poses serious threats to privacy, reputation, democratic institutions, cybersecurity, intellectual property rights, and national security. The proliferation of manipulated digital content has increased incidents of identity theft, financial fraud, misinformation, election manipulation, cyber extortion, defamation, and non-consensual explicit content, exposing significant shortcomings in existing legal frameworks. This paper critically examines the legal implications of deepfakes and synthetic media through a comparative analysis of regulatory approaches adopted in India, the European Union, the United States, China, and other leading jurisdictions. It evaluates existing laws relating to data protection, privacy, intellectual property, cybercrime, intermediary liability, and freedom of expression while highlighting emerging legislative initiatives specifically targeting AI-generated content. The paper also discusses ethical concerns surrounding algorithmic accountability, informed consent, digital authenticity, and platform responsibility. It concludes that combating the risks associated with deepfakes requires comprehensive legal reforms, international cooperation, robust AI governance frameworks, technological detection mechanisms, digital literacy initiatives, and balanced regulation that simultaneously protects innovation, fundamental rights, and democracy.
R. Author· Global Journal of Computing...· 0 citations
Synthetic speech generation and voice-cloning technologies have achieved unprecedented levels of realism, enabling numerous applications in accessibility, virtual assistants, and media production. However, these advancements also introduce significant risks, including identity fraud, impersonation attacks, misinformation, and security breaches. This paper proposes a multimodal fusion framework for synthetic voice detection that combines handcrafted acoustic features with deep spectrogram representations to improve detection robustness and generalization. The proposed architecture employs a Convolutional Neural Network–Bidirectional Long ShortTerm Memory (CNN-BiLSTM) network to capture both spectral artifacts and temporal inconsistencies characteristic of AI-generated speech. To enhance transparency and interpretability, an explainability module incorporating attention visualization and feature attribution techniques is integrated into the detection pipeline. Furthermore, the framework is deployed through a real-time inference interface, demonstrating its practical applicability in cybersecurity, digital forensics, and media authentication scenarios. The findings highlight the effectiveness of combining deep learning, multimodal feature fusion, and explainable artificial intelligence to address the growing challenge of synthetic speech detection.
Mahima Bg, Pallavi Gb· International journal of res...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.