Skip to content

Leveraging computer vision and natural language processing for efficient metadata extraction in digitized books

Aug 2026 · International Journal of Pervasive Computing and Communications · pp. 1-25 · 0 citations · 34 references

TL;DR

A scalable framework that combines page classification, object detection, OCR, NER and rule-based refinement to reduce manual metadata extraction in digital library management is proposed.

Abstract

High-quality book metadata improves digital libraries by supporting cataloging, searchability and automated classification. However, metadata generation from scanned books remains challenging due to limited annotated data sets and optical character recognition (OCR) limitations. This study aims to propose a deep learning-based framework for automatic metadata extraction from digitized books by integrating computer vision and natural language processing (NLP). The framework uses MobileNetV2 to classify title pages, table of contents (ToC) pages and content pages. EfficientDet detects metadata-related regions, such as titles and author information. OCR extracts text from these regions, followed by named entity recognition (NER) and regular expressions to refine the extracted metadata. The framework was evaluated using a custom data set of 188 books published between 1800 and 2021, comprising 857 annotated pages. The page classification model achieved 97.16% accuracy. For object detection, the model obtained average precision (AP) and average recall (AR) scores of 71.7% and 42.0% for title pages, 55.1% and 33.4% for ToC pages and 87.9% and 54.4% for content pages, respectively. This study contributes a scalable framework that combines page classification, object detection, OCR, NER and rule-based refinement to reduce manual metadata extraction in digital library management. Future research can expand multilingual data sets, improve robustness to OCR noise and explore end-to-end learning for richer bibliographic metadata extraction.

View source

Similar papers

Preprint Aug 2026

Institutional Books - Visual Elements: An open-source pipeline for extracting, classifying, deduplicating, and captioning visual elements from digital book collections

An open-source end-to-end pipeline for detecting, classifying, deduplicating, and captioning visual elements from historical book collections is introduced and an initial dataset of 22.6 million visual elements extracted from the Institutional Books: Harvard Library dataset is released.

Jimmy Mendez, Matteo Cargnelutti, David Lowry-Duda et al. · 0 citations
Open access Jul 2026

Handwritten Text Extraction and Digitization

Extracting structured information from visually rich documents remains a complex task due to variations in layout, text alignment, and reading order. Traditional methods based on IOB tagging or graph decoding often struggle with irregular text sequences and the computational burden of large relational graphs. This paper introduces a novel anchor-based approach that redefines entity representation and association for structured information extraction. The proposed model, named Hwte, integrates visual and linguistic features through a multi-modal transformer architecture that jointly detects entities and their relationships. A new pre-training objective, Masked Detection Modelling (MDM), is introduced to enhance the model’s ability to predict both textual and spatial information simultaneously. Experimental evaluations on benchmark datasets demonstrate that the proposed method achieves superior accuracy and robustness compared to existing solutions, highlighting its effectiveness for real-world document understanding tasks.

Anbu Lakshmi S, P. R. Raksha, Mohamadi Ghouisya Kousar et al. · 0 citations
Aug 2026

Automated Indexing of Historical Postcards: An End-to-End Approach Combining Image and Text Analysis

An end-to-end approach for automated historical postcard indexing that integrates computer vision and natural language processing techniques is presented and effective integration of multiple AI techniques for automated heritage document analysis is demonstrated.

Matthieu Pélingre, Salvatore Tabbone · 0 citations

CroQS: Cross-modal Query Suggestion for Text-to-Image Retrieval in Image Archives

The novel task of cross-modal query suggestion is introduced, which interactively guides users by suggesting textual refinements based on visual clusters identified in the retrieval results, and the creation of CroQS, a benchmark dataset comprising 50 diverse queries and 295 semantic clusters in generic domain.

Giacomo Pacini, Nicola Messina, Nicola Tonellotto et al. · 0 citations
Review Aug 2026

Reproducible Multimodal Artificial Intelligence Workflow for Historical Digital Archive Discovery.

Historical digital archives are increasingly searchable, but discovery remains limited when optical character recognition errors, visually heterogeneous document types, and sparse metadata are handled separately. This study aimed to develop a reproducible protocol to evaluate whether a conservative multimodal artificial intelligence workflow can improve archival retrieval without replacing archivist review. A corpus of 1,600 digitized archival records from eight document classes was assembled, and a 420-query benchmark was used to compare five retrieval conditions: baseline indexing, optical character recognition correction alone, visual classification alone, metadata enrichment alone, and full multimodal integration. Record-level outcomes included character error rate (CER), word error rate (WER), named-entity recall, document-type classification accuracy, prediction confidence, metadata completeness, and subject-heading match. Query-level outcomes included P@10, R@10, nDCG@10, time to first relevant result, and successful-search rate. Optical character recognition correction reduced WER most strongly for handwritten letters, ledgers, and registry books; visual fine-tuning improved classification accuracy most for maps, posters, newspapers, and registry books; and metadata enrichment increased completeness across all historical periods. The full multimodal condition achieved the highest retrieval performance (P@10 = 0.624, R@10 = 0.492, nDCG@10 = 0.634) and reduced the mean time to first relevant result from 156.079 s to 58.485 s. These results support a modular, auditable workflow in which text, image, and descriptive signals are combined under explicit thresholds and human review.

Weiwei Zhang, Junmei Gai · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.