NP-SEE: Structural Analysis of Historical and Undeciphered Writing Systems — A Non-Parametric Information-Theoretic Approach
Abstract
Version 3.0 We applied the NP-SEE (Non-Parametric Structural Entropy Engine) v6.1 framework to eleven corpora of historical and literary symbolic representations: Voynich Manuscript (ZL3b + GC2a), Linear A, Linear B, Inca Quipu (Open Khipu Repository, 619 khipu, 54,403 cords), Proto-Elamite, Etruscan, Meroitic, Indus Valley script, Homeric tradition (Iliad + Odyssey), and Shakespearean corpus (five plays). NP-SEE measures structural complexity on encoded symbolic sequences without dictionaries, grammars, or prior linguistic knowledge, using 1,000 bootstrap permutations (Phipson & Smyth 2010). Main results: (1) Voynich is structurally isolated — M4=1 (single global rule), SEG_lambda_std=0.036–0.044, MICRA ROBUST — a profile incompatible with all natural languages, positional systems, and administrative registers in the corpus; (2) Linear A shows the highest structural variability among linguistic systems (Lambda=0.602, SEG_lambda_std=0.058), with 57 anomalous segments out of 166; (3) Proto-Elamite retains the highest chi²=138,228, consistent with an archaic administrative system with fixed categories; (4) Quipu shows the highest Lambda (0.669) and T_index (0.549) with chi²=0, consistent with a strongly position-dependent encoding; (5) Iliad and Odyssey show structurally negligible differences across all SEG indicators (Cohen's d=0.031 on SEG_lambda_mean), a structural observation relevant to the Homeric Question; (6) within the Shakespearean corpus, plays of consolidated attribution and plays recorded as doubtful show a systematic difference in lexical richness (Cohen's d=0.683) and a smaller difference in SEG_lambda_mean (d=0.240). Version 3.0 introduces three new segment-level indicators (SEG_lambda_mean, SEG_lambda_std, SEG_events) and two new literary corpora (Homeric tradition, Shakespearean corpus) not present in v2.1. All thresholds are pre-defined on the basis of v2.1 calibration. Results are exploratory and hypothesis-generating; philological interpretation requires domain expertise. Supersedes v2.1 (DOI: 10.5281/zenodo.22239139). Document written with AI assistance (Claude, Anthropic). Version 2.0 — Major revision Changes from v1.0: Reformulated as a methodological white paper / technical report presenting preliminary results and research hypotheses, not definitive conclusions All philological conclusions rewritten in cautious exploratory language ("indicates", "is compatible with", "does not detect evidence of") Corrected M32 p-values using add-one correction: p=1/51≈0.020 (50 permutations), avoiding statistically improper p=0 reporting Added formal definitions section for all indicators (Lambda, MICRA, Trap Score, chi² HVG, S_disc, M4, M4v2) Added explicit methodological limitations section: non-reproducibility, encoding dependency, multiple testing without Bonferroni correction Added "Input type" column distinguishing textual corpora, transliterations, physically encoded systems (Quipu), and synthetic distributions (Proto-Elamite) Re-analysed Linear B without CSV format delimiters: chi²=820.76 vs 5.051 — the separator was a format artefact, not linguistic structure; dominant bigrams now show real Mycenaean syllables (a-, o, e-) Revised title to reflect methodological-exploratory nature Processed supplementary data uploaded: M4v2 zone-by-zone analysis for all corpora, M32 pairwise results, summary statistics Supplementary materials included: 7 M4v2 zone analysis files (one per corpus) manoscritti_summary_v6.csv — Lambda, MICRA, Trap, chi² for all 9 corpora m32_filo_results.csv — M32 structural comparison for all 13 selected pairs Independent analysis of additional corpora available on request. This report presents the application of NP-SEE (Non-Parametric Structural Entropy Engine) to nine historical and undeciphered writing systems: Voynich Manuscript (two independent transcriptions), Linear A, Linear B, Inca Quipu, Proto-Elamite, Etruscan, Meroitic, and Indus Valley script. NP-SEE is a mathematical framework for structural analysis of symbolic sequences based on MDL theory (Minimum Description Length, Rissanen 1978) and non-parametric bootstrap (1,000 permutations, Phipson & Smyth 2010). The system operates without any prior linguistic knowledge, dictionaries, grammars, or training data. Key results: (1) The Voynich Manuscript is structurally isolated — M4=1 rule across 12 identical zones, not compatible with the tested natural-language structural models included in the corpus; (2) Linear A has significantly more intense structure than Linear B (Lambda 0.602 vs 0.517); (3) The Inca Quipu shows the highest structural distance from all linear writing systems; (4) M32 pairwise comparison between Meroitic and Linear A does not support direct structural affinity under the present NP-SEE analysis; (5) Etruscan shows the highest structural variety in the corpus (14 varied M4v2 zones). The report includes complete zone-by-zone M4v2 analysis and M32 MIVAP pairwise comparisons for all 13 pairs. Document prepared with the support of artificial intelligence tools. Note on supplementary files: Input text files are provided for corpora available in transliterated textual format only: Voynich ZL3b and GC2a (IVTFF transcriptions, https://www.voynich.nu), Linear A and Linear B without CSV delimiters (DĀMOS corpus, University of Oslo, https://damos.hf.uio.no). The remaining five corpora are encoded as binary sequences — source data is publicly available from the original repositories: Inca Quipu: Open Khipu Repository — https://khipukamayuq.fas.harvard.edu | Etruscan: Fowler & Wolfe 1965, Internet Archive — https://archive.org/details/MaterialsForTheStudyOfTheEtruscanLanguage | Meroitic: machine-readable corpus (Otten 2025) — https://github.com/Joshua-Otten/Meroitic-Corpus | Proto-Elamite: synthetic Zipf distribution from CDLI — https://cdli.mpiwg-berlin.mpg.de | Indus Valley: corpus insufficient (1KB) — not uploaded. M4v2 zone analysis files are provided for 7 of the 9 corpora. Inca Quipu is excluded because the Quipu physical-attribute encoding is not compatible with symbolic bigram analysis. Indus Valley is excluded due to insufficient corpus size.