Skip to content
Preprint

FindMyText: Robust, Scalable Detection of Text Containment in Large Web-Crawled Corpora

Jul 2026 · 0 citations · 23 references
Computer Science

Abstract

We present FindMyText, an open-source Python package designed to efficiently assess whether a given text appears, in part or in full, within a text corpus. The tool builds on prior techniques for document fingerprinting, but extends them with a novel mechanism to explicitly capture sequences of matching fingerprints. By identifying such chains, the tool can more reliably detect near-verbatim copies of a given text rather than mere textual similarities. This makes FindMyText particularly suited for verifying the presence of copyrighted material in a corpus. Leveraging a distributed, disk-based indexing framework, the system scales to large web-crawled datasets. Using a new benchmark for evaluating text containment methods, we show that FindMyText outperforms alternative approaches across three datasets (ArXiv papers, Wikipedia, and generic web content).

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.