Skip to content

Author

Liri Sokol

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#software testing Dataset Open access Sep 2026

Text Recognition Model for Historical Persian Newspapers

Text Recognition Model for Historical Persian Newspapers Summary childhood-in-the-iranian-press-persian-newspaper-recognition-v1.mlmodel recognizes printed, right-to-left Persian text lines from historical newspapers. It must be used after page layout and baseline detection; it is not a full-page layout-analysis model. Framework and compatibility Format: Kraken/Ketos .mlmodel Tested software environment: Kraken 5.3.0 Text direction: horizontal right-to-left Model SHA-256: 9f91820f7ef856d093f7960a19da9085f442eb791ebe3060ecb87c4d446861af Parent model and provenance The model was fine-tuned from Benjamin Kiessling's Printed Persian Base Model Trained on the OpenITI Corpus (2022), DOI https://doi.org/10.5281/zenodo.7051644. The project copy of the parent had MD5 1c322f2defb5f2f96f95c15cc68a5169. The upstream model was trained on the Persian subset of OpenITI printed Arabic-script data and was intended as a base model for further fine-tuning. Its model-repository metadata declares Apache-2.0; see THIRD-PARTY-NOTICES.md. Training data persian-newspaper-recognition-alto-training-data-v1.zip contains 68 ALTO XML files from recognition-v2/, and no source images or non-ALTO files. ALTO transcriptions and coordinates are supplied for provenance, but the referenced page images are excluded. Training procedure The checked-in training procedure used Kraken Ketos recognition training with ALTO input, initialized from persian_base.mlmodel, with right-to-left base direction, image resizing, learning rate 0.0001, and eight threads. See TRAINING_COMMAND.txt. The model metadata records 65 completed training epochs. The exact completed command, random seed, GPU/hardware, split, and original training log are not documented for this release. Evaluation EVALUATION.md reports in-sample character accuracy of 94.93% (CER 5.07%) over the 68 accompanying ALTO files. This is not an estimate of performance on unseen material. Limitations and responsible use Performance may degrade on degraded scans, uncommon or decorative typefaces, mixed Latin/Persian content, numerals, non-textual regions, and page layouts not represented in the annotation material. Compare transcription against page images before historical interpretation or quotation. Citation Sokol, L. 2026. Text Recognition Model for Historical Persian Newspapers (Version 1.0.0) [machine-learning model]. Zenodo. https://doi.org/10.5281/zenodo.22537592

Liri Sokol · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.