Author

Dana Halabi

1 paper indexed here

Fetches their full publication history.

Not the right person? Other researchers publish under this name.

Open access Aug 2026

Evaluating lexical feature extraction for plagiarism detection in Arabic documents

Plagiarism detection is the task of determining whether a document contains parts from other documents by employing different styles of plagiarism, such as copying certain parts and reordering or replacing words with synonyms, without citing the original text owner. This task is important in many applications, and there are two primary types of plagiarism detection methods: external and intrinsic. Plagiarism detection in Arabic documents is challenging because of Arabic’s rich morphological features, lexical variation, and syntactic complexity, which limit the effectiveness of some detection approaches. To address these challenges, this study introduces an external plagiarism detection framework built on an artificial neural network (ANN) model and a lexical feature extraction framework adapted to the linguistic features of Arabic. The proposed framework is evaluated using ExAraPlagDet-2015 benchmark, where a baseline model using support vector machine (SVM) is introduced for comparison. Experimental results demonstrate notable improvements in plagiarism detection performance of the proposed framework compared with SVM and other baseline methods. The proposed framework provides a precision value of 92% and an F-score value of 96%, verifying its effectiveness for Arabic plagiarism detection.

Marwah Alian, Dana Halabi, H. Alshboul · 0 citations