Evaluating Sentence Splitting on 19th-Century Italian Novels: a Comparative Analysis Across Approaches
Abstract
The automatic segmentation of raw text into individual sentences, known as sentence splitting or sentence segmentation, is a fundamental task in text processing. Although it is often considered to be solved in standard domains such as news articles and Wikipedia pages, the performance of the system can vary significantly between different textual genres. This study evaluates eight sentence splitting tools employing rule-based, supervised, semi-supervised, and unsupervised approaches, and additionally tests two Large Language Models in a zero-shot setting, on a corpus of 19th-century Italian novels, namely “I Promessi Sposi”, “I Malavoglia”, “Le avventure di Pinocchio”, and “Cuore”. In addition, we train new sentence splitting models using the Stanza pipeline, creating individual models for each novel as well as a combined model trained on all available data. This work aims to highlight that, although literary texts have received relatively little attention in sentence segmentation research, they offer a rich and promising intersection between Natural Language Processing, Italian linguistics, and the Digital Humanities.