Skip to content
#small language model Open access

NP-SEE: Structural Analysis of Historical and Undeciphered Writing Systems — A Non-Parametric Information-Theoretic Approach

Sep 2026 · Zenodo (CERN European Organization for Nuclear Research)

Abstract

Version 3.0 We applied the NP-SEE (Non-Parametric Structural Entropy Engine) v6.1 framework to eleven corpora of historical and literary symbolic representations: Voynich Manuscript (ZL3b + GC2a), Linear A, Linear B, Inca Quipu (Open Khipu Repository, 619 khipu, 54,403 cords), Proto-Elamite, Etruscan, Meroitic, Indus Valley script, Homeric tradition (Iliad + Odyssey), and Shakespearean corpus (five plays). NP-SEE measures structural complexity on encoded symbolic sequences without dictionaries, grammars, or prior linguistic knowledge, using 1,000 bootstrap permutations (Phipson & Smyth 2010). Main results: (1) Voynich is structurally isolated — M4=1 (single global rule), SEG_lambda_std=0.036–0.044, MICRA ROBUST — a profile incompatible with all natural languages, positional systems, and administrative registers in the corpus; (2) Linear A shows the highest structural variability among linguistic systems (Lambda=0.602, SEG_lambda_std=0.058), with 57 anomalous segments out of 166; (3) Proto-Elamite retains the highest chi²=138,228, consistent with an archaic administrative system with fixed categories; (4) Quipu shows the highest Lambda (0.669) and T_index (0.549) with chi²=0, consistent with a strongly position-dependent encoding; (5) Iliad and Odyssey show structurally negligible differences across all SEG indicators (Cohen's d=0.031 on SEG_lambda_mean), a structural observation relevant to the Homeric Question; (6) within the Shakespearean corpus, plays of consolidated attribution and plays recorded as doubtful show a systematic difference in lexical richness (Cohen's d=0.683) and a smaller difference in SEG_lambda_mean (d=0.240). Version 3.0 introduces three new segment-level indicators (SEG_lambda_mean, SEG_lambda_std, SEG_events) and two new literary corpora (Homeric tradition, Shakespearean corpus) not present in v2.1. All thresholds are pre-defined on the basis of v2.1 calibration. Results are exploratory and hypothesis-generating; philological interpretation requires domain expertise. Supersedes v2.1 (DOI: 10.5281/zenodo.22239139). Document written with AI assistance (Claude, Anthropic). Version 2.0 — Major revision Changes from v1.0: Reformulated as a methodological white paper / technical report presenting preliminary results and research hypotheses, not definitive conclusions All philological conclusions rewritten in cautious exploratory language ("indicates", "is compatible with", "does not detect evidence of") Corrected M32 p-values using add-one correction: p=1/51≈0.020 (50 permutations), avoiding statistically improper p=0 reporting Added formal definitions section for all indicators (Lambda, MICRA, Trap Score, chi² HVG, S_disc, M4, M4v2) Added explicit methodological limitations section: non-reproducibility, encoding dependency, multiple testing without Bonferroni correction Added "Input type" column distinguishing textual corpora, transliterations, physically encoded systems (Quipu), and synthetic distributions (Proto-Elamite) Re-analysed Linear B without CSV format delimiters: chi²=820.76 vs 5.051 — the separator was a format artefact, not linguistic structure; dominant bigrams now show real Mycenaean syllables (a-, o, e-) Revised title to reflect methodological-exploratory nature Processed supplementary data uploaded: M4v2 zone-by-zone analysis for all corpora, M32 pairwise results, summary statistics Supplementary materials included: 7 M4v2 zone analysis files (one per corpus) manoscritti_summary_v6.csv — Lambda, MICRA, Trap, chi² for all 9 corpora m32_filo_results.csv — M32 structural comparison for all 13 selected pairs Independent analysis of additional corpora available on request. This report presents the application of NP-SEE (Non-Parametric Structural Entropy Engine) to nine historical and undeciphered writing systems: Voynich Manuscript (two independent transcriptions), Linear A, Linear B, Inca Quipu, Proto-Elamite, Etruscan, Meroitic, and Indus Valley script. NP-SEE is a mathematical framework for structural analysis of symbolic sequences based on MDL theory (Minimum Description Length, Rissanen 1978) and non-parametric bootstrap (1,000 permutations, Phipson & Smyth 2010). The system operates without any prior linguistic knowledge, dictionaries, grammars, or training data. Key results: (1) The Voynich Manuscript is structurally isolated — M4=1 rule across 12 identical zones, not compatible with the tested natural-language structural models included in the corpus; (2) Linear A has significantly more intense structure than Linear B (Lambda 0.602 vs 0.517); (3) The Inca Quipu shows the highest structural distance from all linear writing systems; (4) M32 pairwise comparison between Meroitic and Linear A does not support direct structural affinity under the present NP-SEE analysis; (5) Etruscan shows the highest structural variety in the corpus (14 varied M4v2 zones). The report includes complete zone-by-zone M4v2 analysis and M32 MIVAP pairwise comparisons for all 13 pairs. Document prepared with the support of artificial intelligence tools. Note on supplementary files: Input text files are provided for corpora available in transliterated textual format only: Voynich ZL3b and GC2a (IVTFF transcriptions, https://www.voynich.nu), Linear A and Linear B without CSV delimiters (DĀMOS corpus, University of Oslo, https://damos.hf.uio.no). The remaining five corpora are encoded as binary sequences — source data is publicly available from the original repositories: Inca Quipu: Open Khipu Repository — https://khipukamayuq.fas.harvard.edu | Etruscan: Fowler & Wolfe 1965, Internet Archive — https://archive.org/details/MaterialsForTheStudyOfTheEtruscanLanguage | Meroitic: machine-readable corpus (Otten 2025) — https://github.com/Joshua-Otten/Meroitic-Corpus | Proto-Elamite: synthetic Zipf distribution from CDLI — https://cdli.mpiwg-berlin.mpg.de | Indus Valley: corpus insufficient (1KB) — not uploaded. M4v2 zone analysis files are provided for 7 of the 9 corpora. Inca Quipu is excluded because the Quipu physical-attribute encoding is not compatible with symbolic bigram analysis. Indus Valley is excluded due to insufficient corpus size.

View source

Similar papers

#small language model Dataset Open access Oct 2026

Socratic guiding questions in synthetic arithmetic data: matched LoRA runs (revision v2)

Supporting data, adapters, predictions and code for the article *Low-Cost LoRA Fine-Tuning of Small Language Models for Multi-Step Arithmetic Reasoning* by Jake O'Grady, Asena Isik Gürhan, Chee Fong Ting and Effirul Ramlan (University of Galway). We generated 20,000 GSM8K-derived arithmetic problems with step-by-step s...

O'Grady, Jake, Gürhan, Asena Isik, Chee, Fong Ting et al. · 465 citations
#computer vision Open access Jun 2016

Software Development in Startup Companies: The Greenfield Startup Model

The results are packaged in the Greenfield Startup Model (GSM), which explains the priority of startups to release the product as quickly as possible, and the need to shorten time-to-market, by speeding up the development through low-precision engineering activities.

Carmine Giardino, Nicolò Paternoster, M. Unterkalmsteiner et al. · 178 citations · ⚡14
#computer vision Open access Oct 2016

Software Startups - A Research Agenda

Software startup companies develop innovative, software-intensive products within limited timeframes and with few resources, searching for sustainable and scalable business models.

M. Unterkalmsteiner, P. Abrahamsson, Xiaofeng Wang et al. · 157 citations · ⚡17
#machine learning Review Open access Oct 2016

“Failures” to be celebrated: an analysis of major pivots of software startups

This study conducts a case survey study based on the secondary data of the major pivots happened in 49 software startups, and demonstrates that customer need pivot is the most common among all pivot types.

Sohaib Shahid Bajwa, Xiaofeng Wang, Anh Nguyen-Duc et al. · 127 citations · ⚡15
#computer vision Review Open access May 2015

A survey study on major technical barriers affecting the decision to adopt cloud services

The comparison of adopter and non-adopter sample reveals three potential adoption inhibitor, security, data privacy, and portability, which underlines the importance of the technical and security perspectives for research investigating the adoption of technology.

Nattakarn Phaphoom, Xiaofeng Wang, S. Samuel et al. · 111 citations · ⚡8
#computer vision Open access Feb 2018

Lean Internal Startups for Software Product Innovation in Large Companies: Enablers and Inhibitors

This study investigates how Lean internal startup facilitates software product innovation in large companies and identifies its enablers and inhibitors, and shows the potential of the method-in-action framework to investigate the Lean startup approach in non-startup context.

Henry Edison, Nina M. Smørsgård, Xiaofeng Wang et al. · 78 citations · ⚡6

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.