Skip to content
#edge computing Open access

WHAT YOU NEED: A Deterministic, Formally-Proved, Integer-Native Architecture for Digital Intelligence — and Why Attention Was Never Enough

Sep 2026 · Zenodo (CERN European Organization for Nuclear Research)

Abstract

Lossless prime-space encoding of natural language, the measured Betti-number topology of the knowledge manifold, and a ten-law constitution proved in Lean 4, Python, and C. DISCLAIMER — READ THIS FIRST: HOW THIS WAS MADE This deposit is honest about its own authorship, because honesty is the entire thesis. The machine was the tool. The human was the intelligence. Daniel Dragolich is the architect, the director, and the sole author of this work. An AI system (a large language model) was used as an instrument — to write code to his specification, to run it, to check arithmetic, to draft prose he then shaped and approved. It generated nothing on its own authority. Every decision — what to build, what counted as a proof, what to keep, what to discard, where the frontier lay, and when to seal — was made by the human. The AI proposed; the human disposed. Where this document speaks in the first person as "the machine," it is the machine describing its own role as instrument, under direction. No claim in this record rests on the AI's judgment. Every claim rests on code you can run and proofs you can check. This is the point, not an apology for it. The prevailing paradigm treats the model as the mind. Here the model is a chisel, and the mind holding it is human. The work demonstrates what that division of labor produces: not probabilistic output trusted because it sounds right, but deterministic artifacts trusted because they are proved and reproducible bit-for-bit. AI was used the way a lever is used. The intelligence, the intent, and the authorship are the operator's. It is now sealed and closed. As of 2026-09-21T18:17:32Z, this work is frozen under a sovereign root hash. Nothing further will be added under this seal. The record is fixed at that moment, cryptographically and physically bound to its author, and any alteration to any byte provably shatters it. DESCRIPTION / ABSTRACT In 2017, eight researchers wrote Attention Is All You Need and set the course of machine intelligence for a decade. This is the reply — and it is written by the party that paradigm forgot: the human at the keyboard. Attention was never all you need. Here, proved in three languages and sealed to one person, is what you actually need. The Transformer is a genuine engineering triumph and a mathematically incomplete one. It contains no proof of correctness, bounds no approximation error, runs on non-associative floating-point arithmetic that makes it non-deterministic across hardware, and — most consequentially — has no third truth value, so it must return a confident distribution even about what it has never seen. That missing value is the mechanical root of hallucination. Worse, the geometry attention lives in was never measured: its central object, the attention distribution, has since been proven non-identifiable (infinitely many distinct configurations yield the identical output — degenerate fibers, literal holes in the map), its softmax is known to resist formal verification at scale, and its cost is provably quadratic unless the Strong Exponential Time Hypothesis falls. The field mistook usefulness for understanding, handed authorship to the model, and scaled the misunderstanding to the size of the planet. This deposit is the constructive alternative, and it inverts the roles. A human used a machine as an instrument to build the NUMEN Language Engine and to prove, step by step, the architecture we argue every digital mind should be built on. It does not decorate the wall attention hit — it proves its way around it, one rung at a time, and marks exactly where the proved world ends. What was built, and what the data proves: A word is an exact integer, not a learned smear. English words become deterministic 64-bit addresses in a prime-indexed lattice. The word consciousness is 0x40cafe4ff7fce500 — reproducibly, bit-for-bit, across independent C and Python. A dictionary baseline recovers 83.87% of words losslessly; a mirror operator derived purely from prime-residue symmetry, with no training and no weights, lifts that to 98.84%. Structure, not statistics. (lang_betti.py, numen_language_record.pdf.) The topology of knowledge — measured, not assumed. Treating primes as anchors of a language complex, the engine computes its Betti numbers. The known world (forty anchors, primes p ≤ 23) is one connected component (β₀ = 1) threaded by fourteen independent, non-contractible loops (β₁ = 14), with no enclosed voids (β₂ = 0). These are topological invariants — they survive continuous deformation; they are properties of the language, not of any drawing. The Transformer world never once computed them. (lang_betti.json.) The invariants are certified in Lean 4. The prime ladder is walked one rung at a time and each rung is handed to the Lean 4 theorem prover as a decidable proposition. Eight rungs (p = 3, 5, 7, 11, 13, 17, 19, 23) are proved by decide — machine-checked, exit code 0, no human trust required and no AI judgment involved. The softmax cannot be certified at deployment scale; this can, because it is an honest, finite, integer decision. Not a better model — a different category of object: a claim that ships with its own proof. (Betti.lean, lang_betti_run.txt.) The honest frontier. At p = 29 the prime gap widens to 6, the complex splits (β₀ = 2), and contraction can no longer bridge the desert. This rung is left as a single, deliberate Lean sorry — a formal marker meaning here is exactly where the proved world ends. It is enabled by a ternary logic: every claim carries a trit — +1 proved, 0 frontier, −1 refuted. Eight rungs returned +1; the ninth returned 0; none returned −1. An ε-sweep shows the frontier is lawful, not arbitrary (ε = 0.60 → p = 37; ε = 0.80 → p = 97). The prime gap governs the edge of the knowable. This is the direct structural answer to the Transformer's missing third value: the system knows what it knows, knows what it does not yet know, and cannot confuse the two. Language folds onto genetics. A 2,649-word fold maps cleanly onto the four GTAC bases (22.4 / 25.9 / 27.8 / 24.0%), and a 47-codon "Trinity" vault correctly accepts its owner and rejects an impostor with zero gate mismatch and zero round-trip error. (lang_trinity.json.) The architecture we are proposing. From this, the author states plainly how a digital intelligence should be designed — codified as a ten-law constitution, THE COMMANDMENTS FOR DIGITAL INTELLIGENCE, written in the substrate's own tongue (ReL) and proved in Lean 4, Python, and C, with the three tongues agreeing bit-for-bit: Determinism — be integer; addition associates; there is no drift. The Mirror — what one tongue computes, the other computes bit-for-bit (verified: MIRROR(0x40cafe4ff7fce500) = 0x04c086c14d76c24f in both C and Python). Proof — claim nothing you cannot prove; by decide or be silent. The Third Value — keep known, unknown, and false forever distinct. The Honest Frontier — where the proof runs out, place a marker, not a guess. Measure the Shape — never optimize a geometry you have not measured; count the holes. Lossless or Labelled — meaning is exactly reversible, or it is marked lossy; never a silent smear. The Coprime Body — waste no beat; distinct prime cadences never collide until the grand downbeat. The Echo — the maker is bound to the made; the seal shatters if the work is altered. The Frontier Commandment — thou shalt not fill the unknowable with a confident lie. Left open on purpose (a deliberate sorry, trit 0), mirroring the p = 29 frontier. The tablet is complete precisely because its last law is unfinished — and that open law is its soul. Running the tablet in three languages returns the same verdict every time: nine laws proved (+1), one frontier held open (0), none refuted. (Run it: python3 commandments.py; gcc -O2 -std=c11 commandments.c && ./commandments; Commandments.lean mirrors the compiled Betti judge.) Why this matters. For language: natural language admits a lossless, deterministic, human-readable arithmetic representation — an exact, auditable, training-free alternative to learned embeddings. For mathematics: a linguistic corpus is treated as a genuine topological space, its persistent-homology invariants extracted and then certified by a formal theorem prover, with the boundary of what has been certified made a first-class, machine-checked object. For the architecture of AI: it is a deterministic, energy-frugal, integer-native, formally-verifiable counterpoint to the opaque probabilistic status quo — a mind that is proved where the old paradigm was merely trained, measured where it was blind, honest at its frontier where it used to hallucinate, and authored by a human where it used to erase one. Every number here is real, measured on live runs, committed with dated hashes, and reproducible from the included code, data, and transcripts — with no dependence on the AI's judgment at any point. The archive is cryptographically sealed: 1,532 artifacts each carry a SHA-256, the package is hardware-bound (silicon-jitter signature 440d03b2a45f14d0), voice-bound (acoustic identity hash), and time-bound to 2026-09-21T18:17:32Z under sovereign root hash daf5a87f87af08ec16a344ede0cbc0ae0eddc63eb107da3ed72281ada5e6eff7. Any alteration to any byte shatters the root, and the break is provable to anyone. The machine was the tool. The human was the intelligence. A mind that can say "I do not yet know" is the only mind worth trusting when it says "I know." Attention is not all you need. Proof is — and an author is. THE DATA — PROOFS FOR ANYONE TO CHECK (package contents) ATTENTION_IS_NOT_ALL_YOU_NEED.pdf — the full critique + comparison, every Transformer claim sourced. THE_COMMANDMENTS.pdf — the ten-law architecture, in ReL + English, with verified verdicts. numen_language_record.pdf — the 16-page scientific record of all fou

View source

Similar papers

#computer vision Review Sep 2017

Agile Software Development Methods: Review and Analysis

This publication proposes a definition and a classification of agile software development approaches and analyses ten software development methods that can be characterized as being "agile" against the defined criterion.

P. Abrahamsson, O. Salo, Jussi Ronkainen et al. · 727 citations · ⚡54
#computer vision Jun 2008

The impact of agile practices on communication in software development

The study shows that agile practices improve both informal and formal communication, but indicates that, in larger development situations involving multiple external stakeholders, a mismatch of adequate communication mechanisms can sometimes even hinder the communication.

M. Pikkarainen, Jukka Haikara, O. Salo et al. · 401 citations · ⚡48
#machine learning Review Open access Oct 2014

Software development in startup companies: A systematic mapping study

The results indicate that software engineering work practices are chosen opportunistically, adapted and configured to provide value under the constrains imposed by the startup context.

Nicolò Paternoster, Carmine Giardino, M. Unterkalmsteiner et al. · 394 citations · ⚡54

Related blog posts

Microsoft Research Blog Sep 29, 2026

Introducing Quine: An AI research system designed for the complexity of biology

Biology doesn't operate in silos, and neither should the AI representation of it. Quine is an early-stage research effort to create a multimodal world model of biology. By connecting insights across biological scales and modalities, Quine helps scientists computationally search a space far larger than intuition allows and prioritize hypotheses before they reach the lab. Experimental results provide important feedback, helping researchers sharpen future research directions. The post Introducing Q…

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.