Aug 2026· International Journal of Pervasive Computing and Communications· pp. 1-25· 0 citations· 34 references
TL;DR
A scalable framework that combines page classification, object detection, OCR, NER and rule-based refinement to reduce manual metadata extraction in digital library management is proposed.
Abstract
High-quality book metadata improves digital libraries by supporting cataloging, searchability and automated classification. However, metadata generation from scanned books remains challenging due to limited annotated data sets and optical character recognition (OCR) limitations. This study aims to propose a deep learning-based framework for automatic metadata extraction from digitized books by integrating computer vision and natural language processing (NLP).
The framework uses MobileNetV2 to classify title pages, table of contents (ToC) pages and content pages. EfficientDet detects metadata-related regions, such as titles and author information. OCR extracts text from these regions, followed by named entity recognition (NER) and regular expressions to refine the extracted metadata. The framework was evaluated using a custom data set of 188 books published between 1800 and 2021, comprising 857 annotated pages.
The page classification model achieved 97.16% accuracy. For object detection, the model obtained average precision (AP) and average recall (AR) scores of 71.7% and 42.0% for title pages, 55.1% and 33.4% for ToC pages and 87.9% and 54.4% for content pages, respectively.
This study contributes a scalable framework that combines page classification, object detection, OCR, NER and rule-based refinement to reduce manual metadata extraction in digital library management. Future research can expand multilingual data sets, improve robustness to OCR noise and explore end-to-end learning for richer bibliographic metadata extraction.
An open-source end-to-end pipeline for detecting, classifying, deduplicating, and captioning visual elements from historical book collections is introduced and an initial dataset of 22.6 million visual elements extracted from the Institutional Books: Harvard Library dataset is released.
Jimmy Mendez, Matteo Cargnelutti, David Lowry-Duda et al.· 0 citations
Extracting structured information from visually rich documents remains a complex task due to variations in layout, text alignment, and reading order. Traditional methods based on IOB tagging or graph decoding often struggle with irregular text sequences and the computational burden of large relational graphs. This paper introduces a novel anchor-based approach that redefines entity representation and association for structured information extraction. The proposed model, named Hwte, integrates visual and linguistic features through a multi-modal transformer architecture that jointly detects entities and their relationships. A new pre-training objective, Masked Detection Modelling (MDM), is introduced to enhance the model’s ability to predict both textual and spatial information simultaneously. Experimental evaluations on benchmark datasets demonstrate that the proposed method achieves superior accuracy and robustness compared to existing solutions, highlighting its effectiveness for real-world document understanding tasks.
Anbu Lakshmi S, P. R. Raksha, Mohamadi Ghouisya Kousar et al.· International Research Journ...· 0 citations
An end-to-end approach for automated historical postcard indexing that integrates computer vision and natural language processing techniques is presented and effective integration of multiple AI techniques for automated heritage document analysis is demonstrated.
Matthieu Pélingre, Salvatore Tabbone· Journal on Computing and Cul...· 0 citations
The novel task of cross-modal query suggestion is introduced, which interactively guides users by suggesting textual refinements based on visual clusters identified in the retrieval results, and the creation of CroQS, a benchmark dataset comprising 50 diverse queries and 295 semantic clusters in generic domain.
Giacomo Pacini, Nicola Messina, Nicola Tonellotto et al.· 0 citations
Historical digital archives are increasingly searchable, but discovery remains limited when optical character recognition errors, visually heterogeneous document types, and sparse metadata are handled separately. This study aimed to develop a reproducible protocol to evaluate whether a conservative multimodal artificial intelligence workflow can improve archival retrieval without replacing archivist review. A corpus of 1,600 digitized archival records from eight document classes was assembled, and a 420-query benchmark was used to compare five retrieval conditions: baseline indexing, optical character recognition correction alone, visual classification alone, metadata enrichment alone, and full multimodal integration. Record-level outcomes included character error rate (CER), word error rate (WER), named-entity recall, document-type classification accuracy, prediction confidence, metadata completeness, and subject-heading match. Query-level outcomes included P@10, R@10, nDCG@10, time to first relevant result, and successful-search rate. Optical character recognition correction reduced WER most strongly for handwritten letters, ledgers, and registry books; visual fine-tuning improved classification accuracy most for maps, posters, newspapers, and registry books; and metadata enrichment increased completeness across all historical periods. The full multimodal condition achieved the highest retrieval performance (P@10 = 0.624, R@10 = 0.492, nDCG@10 = 0.634) and reduced the mean time to first relevant result from 156.079 s to 58.485 s. These results support a modular, auditable workflow in which text, image, and descriptive signals are combined under explicit thresholds and human review.
Weiwei Zhang, Junmei Gai· Journal of Visualized Experi...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.