Skip to content
#edge computing Open access

The Letter-"a" sequence from EOA Encoding E5 at LCR = 3.14159: Visualizer, Reproducible Report, and Open Shift-Space Challenge

Sep 2026 · Zenodo (CERN European Organization for Nuclear Research)

Abstract

This deposit contains one finite symbolic word, the tool used to analyse it, and the report the tool produces. The word was made by applying the EOA beta-operator to a phonetic encoding called E5. The input was the spoken English letter name "a". The control parameter, LCR, was set to 3.14159. The word itself, its SHA-256 hash, and every statistic drawn from it are public. The beta-operator, the exact content of encoding E5, and the formal definition of LCR are available under a signed collaboration agreement. This is a single word deposit. It is not a corpus. It is not a dataset. That is on purpose. One word, fully specified, with an open question attached. The tool The tool is a single HTML file. It runs in a browser. No installation, no server, no build step. You just paste your sequence. It produces twelve visual panels and a report. The report embeds the SHA-256 of the word, the definition of every model used, and splits every reported number into three groups: what is observed directly in the word, what is estimated from a chosen model, and what remains an open question. Why the tool is built this way A finite word lets you compute many things. A finite word does not by itself define a shift space. Every number the tool reports is a property of a named object: the word, or a graph built from its factors. No number is claimed as a property of some underlying system. Three objects appear and stay separate throughout: M_k^stride: a letter level stride matrix, M[a,b] = count of positions i where w_i = a and w_(i+k) = b. It is used only to build the projection plane in Panel 5. G_k: the order k factor overlap graph. Nodes are distinct length k factors of the word. Edges are observed overlaps between them. It is used for Panel 6 and Panel 10. A_k: the binary adjacency matrix of G_k. Edge weights are ignored for the spectral calculation. This is the standard object for a topological entropy style quantity on an allowed transition graph. The reported h_k = ln rho(A_k) is the topological entropy of the sofic shift presented by the empirical graph G_k. For that object it is a well defined number, by a standard theorem. It is not claimed to be the topological entropy of any underlying system, since no such system is defined here. The twelve panels Symbolic Turtle Rendering. Maps letters to movement and turning rules and draws the path. The point is to see whether fractal structure comes from the word or from the mapping rule, so the turning strength is adjustable. Chaos Game Representation. Shows which symbol transitions actually occur and how often. Empty regions are transitions that never happen in the word. Recurrence Matrix. Marks repeated symbol positions. Diagonal lines mean periodic structure. Blocks mean repeated subwords. Rewrite Engine. Applies a morphism that you supply, letter by letter, and renders the resulting word as a waveform. If no morphism is given, it draws the raw symbol sequence. Empirical Eigenvector Projection. Walks the prefix Parikh vector and projects it onto a plane built from the dominant and second eigenvector directions of the stride matrix M_k^stride. This is a visual of the empirical projection. It is not a substitution Rauzy fractal. Subword Transition Graph. The order k factor overlap graph G_k, drawn with nodes as distinct k length factors and edges as observed overlaps. It reports node count, edge count, and h_k = ln rho(A_k) at the chosen order. Desubstitution Test. Tries to parse the word under a morphism you supply. It reports a unique parse, an ambiguous parse, or no parse at all. A failed parse is not a proof that the word is non morphic. Factor Complexity p(k) vs k. Number of distinct length k subwords, plotted on log log axes, with Sturmian and linear reference lines. Abelian Complexity a(k) vs k. Number of distinct Parikh vectors of length k factors. Finite Sample Block Growth Ladder. Plots h_k = ln rho(A_k) for each order k, and the finite sample block growth rate (1/n) ln p(n), on the same axes. The gap between them shows how tight the empirical graph is. Prefix Frequency Discrepancy. Tracks the maximum deviation between prefix symbol counts and the full word frequency vector, over the prefix. Surrogate Controls. Compares the original CGR against three controls: a frequency shuffle, a Markov order 1 surrogate, and a periodic control. The controls are seeded from the word, so the same word always produces the same panel. The open question The word shows slow factor complexity growth, near flat abelian complexity at low orders, and a bounded prefix frequency discrepancy. Those features fit several different structural regimes. They do not prove any of them. The question this deposit asks is: given the finite word w produced by the beta-operator from encoding E5 of letter "a" at LCR = 3.14159, which family of shift spaces X has w as a factor of some element, and under what conditions? Possible constructions include a limit sequence, an orbit closure, a substitution subshift, an S-adic system, a language shift, or an empirical finite word system. These are different objects and they are not equivalent. The tool does not pick one. It is built so a specialist can bring their own methods to the released word and to the graphs it defines. What is public and what is not Public: the word, its SHA-256, all finite word statistics, the tool source, and the report. Not public: the beta-operator, the exact content of encoding E5, and the formal definition of LCR. Those are available under a signed collaboration agreement. So the deposit is reproducible at the level of the released word. It is not reproducible from the semantic input alone. Both facts are stated in the report and both are intended. Who this is for Symbolic dynamicists, combinatorics on words researchers, and anyone working on finite word inference of shift space properties. The tool does not argue for a conclusion, but it measures. Call for collaboration If you work on shift spaces, sofic presentations, substitutive systems, or S-adic constructions, and you want to test your methods on a word whose structure is not known to us either, get in touch. The beta-operator, encoding E5, and the EOA technical documentation are available under a standard agreement covering use, attribution, and non redistribution. Two concrete forms of collaboration: Adversarial analysis. Bring your own tests. If you can show the word is incompatible with a proposed regime, or compatible with a stronger claim than we have made, the finding gets published alongside the tool. Definition of the underlying shift space. If you can propose a shift space X that is naturally motivated by the operator and admits a proof of any property we currently leave open, that is the most useful contribution this deposit could receive. Contact: https://research.neurolabs.space/ Version This is v4.1. The tool, the report format, and the definitions are stable. A statistical extension with replicate surrogate ensembles and multiple testing control is planned but not implemented, and is not claimed here.

View source

Similar papers

#computer vision Review Sep 2017

Agile Software Development Methods: Review and Analysis

This publication proposes a definition and a classification of agile software development approaches and analyses ten software development methods that can be characterized as being "agile" against the defined criterion.

P. Abrahamsson, O. Salo, Jussi Ronkainen et al. · 727 citations · ⚡54
#computer vision Jun 2008

The impact of agile practices on communication in software development

The study shows that agile practices improve both informal and formal communication, but indicates that, in larger development situations involving multiple external stakeholders, a mismatch of adequate communication mechanisms can sometimes even hinder the communication.

M. Pikkarainen, Jukka Haikara, O. Salo et al. · 401 citations · ⚡48
#machine learning Review Open access Oct 2014

Software development in startup companies: A systematic mapping study

The results indicate that software engineering work practices are chosen opportunistically, adapted and configured to provide value under the constrains imposed by the startup context.

Nicolò Paternoster, Carmine Giardino, M. Unterkalmsteiner et al. · 394 citations · ⚡54

Related blog posts

Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.