Skip to content
Open access

Detecting Plagiarized Text in Images using OCR and NLP-Based Deep Learning Approaches

Jul 2026 · International journal of advances in soft computing and its applications · Vol 18, pp. 376-384 · 0 citations · 24 references

TL;DR

A novel model, Detecting Embedded Plagiarized Text in Images (DEPTI), is introduced to identify plagiarized text embedded within images, demonstrating high accuracy and robust performance.

Abstract

This paper evaluates plagiarism detection using deep learning and natural language processing (NLP) techniques. A novel model, Detecting Embedded Plagiarized Text in Images (DEPTI), is introduced to identify plagiarized text embedded within images, demonstrating high accuracy and robust performance. DEPTI effectively recognizes paraphrased, translated, and artificial intelligence generated content, achieving strong detection capabilities across diverse scenarios. The model integrates PAN-PC-11, TF-IDF, Tesseract OCR, DistilBERT, and LSTM to extract and analyze text from images, enabling advanced plagiarism detection beyond conventional approaches. Experimental results confirm DEPTI’s effectiveness, highlighting its potential as a reliable tool for safeguarding academic integrity in the digital era.

Read PDF

Similar papers

Open access Aug 2026

Arabic Plagiarism Detection Using Word2Vec-Based Semantic Features and Random Forest Classification on the ExAraPlagDet Dataset

The findings underscore the potential of advanced NLP techniques to overcome language-specific challenges, providing a foundation for future research in multilingual plagiarism detection and enhancing the development of tools for other languages facing similar challenges.

Hanan Fawzy, Ahmad Salah, Heba El-Fiqi et al. · 0 citations
Open access 2026

Combining probabilistic features and semantic features for AI-Generated text detection

The proliferation of Large Language Models (LLMs) such as ChatGPT and Gemini has resulted in a surge of AI-generated text across various domains. However, the widespread use of this technology raises concerns regarding the generation of misinformation and malicious content. To address this challenge, we propose a novel AI-generated Text Detection model combining Probabilistic and Semantic features (ATDPS). Our model extracts semantic features using a pre-trained language model and combines them with probabilistic features generated by multiple LLMs. A temporal convolutional network is employed to process sequence probabilistic features, effectively capturing temporal characteristics within the text. To ensure data coherence and diversity, our dataset includes text generated by a variety of LLMs, including the latest models like GPT-4. Experimental results demonstrate ATDPS's superior performance over existing baselines in terms of accuracy, precision, recall and F1 score, highlighting its potential and effectiveness in detecting AI-generated text.

Yang Yu, Wang Gao · 0 citations

VaryBalance: Detecting LLM-generated Text through Variation

The core of VaryBalance is that, compared to LLM-generated texts, there is a greater difference between human texts and their rewritten version via LLMs, and quantifies this through Mean Squared Deviation and distinguishes human texts and LLM-generated texts.

Xuecong Li, Xiaohong Li, Qiang Hu et al. · 0 citations
Open access Aug 2026

Evaluating lexical feature extraction for plagiarism detection in Arabic documents

This study introduces an external plagiarism detection framework built on an artificial neural network model and a lexical feature extraction framework adapted to the linguistic features of Arabic, verifying its effectiveness for Arabic plagiarism detection.

Marwah Alian, Dana Halabi, H. Alshboul · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.