Skip to content
Open access

Advancing Arabic Text Detection: A CNN-Heuristic and EfficientNet-B1-Based Framework

Aug 2026 · Iraqi Journal of Science · 0 citations · 23 references

Abstract

     In computer vision, word detection in images is still a major challenge, especially for applications like scene interpretation, document indexing, and machine translation. Arabic script presents particular difficulties because of its cursive style, the use of diacritical marks, and the range of typefaces and orientations, which sometimes lead to errors in standard models, even though this effort has achieved great strides in Latin scripts. In this work, we propose a two-step method for Arabic Text Detection (ATD), aiming to improve accuracy without overly complex pipelines. Our approach starts with a set of heuristic rules to pre-select candidate regions. These rules rely on basic geometric and statistical cues, which helped us eliminate much of the irrelevant background noise in initial tests. To further improve the choices, we use an EfficientNet-B1 CNN.  This model uses Global Average Pooling and a final Softmax layer to distinguish between text and non-text regions.  The evaluation, which was conducted on a broad dataset of Arabic scene photos, revealed that combining heuristics with CNN verification enhances results significantly. For instance, we observed an increase in detection accuracy from 89.45% (heuristics alone) to 96.46% with the full pipeline. Moreover, the CNN classifier itself reached an accuracy of 97.77%. These findings confirm the robustness of our approach, particularly in cluttered visual environments.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.