Skip to content
Preprint

Mapping Armenian Paris: Extracting and Geocoding Commercial Advertisements from the 20th-Century Diaspora Press

Aug 2026 · 0 citations · 19 references
Computer Science

TL;DR

An end-to-end, IIIF-based pipeline that turns the digitised Armenian press of France into an interactive map of the 20th-century Parisian Armenian commercial community is presented, showing that VLM-driven data bootstrapping is an effective lever for under-resourced historical languages such as (Western) Armenian.

Abstract

This paper presents an end-to-end, IIIF-based pipeline that turns the digitised Armenian press of France into an interactive map of the 20th-century Parisian Armenian commercial community. On each page, commercial advertisements are located, read, and parsed into structured records, which are then geocoded and placed on the map. Western Armenian is under-resourced and unsupported by off-the-shelf layout and OCR models, so the pipeline uses vision-language models (VLMs) as a data-bootstrapping strategy: they produce usable structured records at a scale hand annotation could not reach, and stay reliable on the strongly curved scans where conventional line-level CRNN OCR breaks down. The contribution includes a 500-page Western Armenian press corpus with 3,270 advertisement-level annotations, a Label Studio template that captures detection and semantic fields in a single annotation pass, and a reproducible workflow transposable to other under-resourced historical corpora. More broadly, the work shows that VLM-driven data bootstrapping is an effective lever for under-resourced historical languages such as (Western) Armenian.

View source

Similar papers

Aug 2026

Automated Indexing of Historical Postcards: An End-to-End Approach Combining Image and Text Analysis

An end-to-end approach for automated historical postcard indexing that integrates computer vision and natural language processing techniques is presented and effective integration of multiple AI techniques for automated heritage document analysis is demonstrated.

Matthieu Pélingre, Salvatore Tabbone · 0 citations
Preprint Aug 2026

OCR-Based Field Extraction for Archaeological Pottery Metadata: The CENTURIA Dataset

Pottery is a primary source for reconstructing the chronological and economic dimensions of past societies. Archaeologists often document ceramic finds through technical drawings and handwritten metadata. This metadata is critical for dating, provenance attribution, and cross-site comparison, but remains inaccessible to computational analysis, requiring manual transcription of every record. We investigate whether state-of-the-art document analysis models can address this task, and introduce CENTURIA, a dataset of 507 pottery records from the Roman site of Carnuntum, providing transcriptions, bounding boxes, and structured field-level labels across seven metadata categories. Benchmarking five OCR models reveals a substantial domain gap: zero-shot transcription error reaches 15-32% SpACER-M, far exceeding rates on printed archival documents, with domain-specific fields recovered in fewer than 3% of cases. LoRA fine-tuning on just 57 samples, reflecting a realistic archival annotation budget, closes this gap, reducing transcription error to below 1.5% and recovering overall field-level accuracy above 87%. Our results show that a small expert-validated fine-tuning set suffices to convert handwritten pottery documentation into structured, searchable metadata ready for archaeological databases.

Gissu Valentina Naghavi, Dominik Hagmann, M. Kampel et al. · 0 citations
Preprint Aug 2026

Institutional Books - Enriched Text: A customizable multilingual open-source pipeline for denoising, deduplicating, and annotating OCR text at scale

Released in 2025, Institutional Books: Harvard Library (IB-HL) is a collection of 983,004 volumes (242B o200k_base tokens), originally digitized through Harvard Library's participation in the Google Books Library project. As researchers and developers have begun to use IB-HL, a tension has emerged between standard large-scale preprocessing practices and the goals of careful information stewardship. Many existing pipelines optimize for web text: as a result, they tend to aggressively filter, deduplicate, restrict by language, and sometimes discard meaningful metadata. Meanwhile, researchers seeking to use IB-HL duplicate effort while performing similar processing and analysis. We describe an approach that we call Enriched Text. Instead of producing a single'complete'stream of tokens, we normalize the text while preserving metadata through annotations. We separate endmatter, detect per-paragraph language, identify clusters of duplicate paragraphs, and compute per-paragraph bits-per-byte scores. We provide this information through HTML-like annotations layered on top of the text. By parsing these annotations, users can tailor the output to their own needs instead of accepting a global editorial decision on content. The pipeline applies to all $\approx$250 languages in the collection. This report describes this project's goals, implementation, and design rationale. The release includes IB-HL-ET (an enriched-text version of IB-HL containing 217B o200k_base tokens across 983,003 volumes, organized into 1.39B annotated subtopic paragraphs) and the pipeline that produced it. These serve to make the collection easier for machines to parse and for humans to study.

David Lowry-Duda, Matteo Cargnelutti, Catherine Brobston et al. · 0 citations
Review Open access Aug 2026

Spatiotemporal Analysis of OpenStreetMap Editing Activities in Japan Using the OSMCha

Abstract. OpenStreetMap (OSM) is a prominent platform for Volunteered Geographic Information (VGI), wherein geographic data are updated daily by a global community of contributors. The OpenStreetMap Changeset Analyzer (OSMCha) serves as a quality assurance tool that automatically flags suspicious edits based on detection rules covering geometric and tag plausibility, edit scale, and contributor behavioral patterns. In this study, we developed Python scripts to systematically collect 740,038 changesets spanning four years (2022–2025) for Japan via the OSMCha API, consolidizing them into a FlatGeobuf database with 60 reason_id mappings embedded as attributes and prefecture-level spatial joins applied. Our analysis revealed a 63.3% increase in annual changesets and a 72.5% expansion in unique contributors, alongside an 11.9- percentage-point decline in the suspicious edit rate. Mobile editors and survey-based edits grew rapidly, consistently demonstrating lower suspicion rates than non-survey edits. A structural shift in the OSMCha detection logic was empirically identified in 2024. The KDE analysis confirmed editing hotspots in three major metropolitan areas, while the January 2024 Noto Peninsula earthquake triggered concentrated crisis mapping, engaging 1,536 contributors.

Toshikazu Seto · 0 citations
#small language model Preprint Aug 2026

Institutional Newspapers Pipeline: Deriving billions of high quality tokens from historical newspapers

The Institutional Newspapers Pipeline is presented, a modular system designed to extract high-quality, structured datasets from historical newspaper scans that was architected so that each step remains interpretable and customizable, and so that the pipeline as a whole remains computationally frugal enough to run on workstation-level hardware.

Matteo Cargnelutti, Catherine Brobston, Eben English et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.