Skip to content
#protein folding Dataset Open access

Touchstone

Sep 2026 · Zenodo (CERN European Organization for Nuclear Research)

Abstract

Touchstone alpha v1.0 What it is Four HTML files that open in any browser, offline, sending nothing anywhere. Plus two Python twins of the same engines for terminal or Colab work. Copyright © 2026 David Coates, MIT. A touchstone is a piece of dark stone: rub gold against it and the streak tells you what the metal really is. These files do that to a number. Who it's for Anyone holding a number and wondering whether it means anything. That's you specifically — working alone, without a supervisor, on patterns nobody has checked. But it's equally a teaching instrument, because the honest answer is usually no, and it shows you why. The one rule everything follows A number is worth something only once it has predicted something it hadn't already seen. The sequence predictor counts only terms a rule got right after the data pinned the rule down. The relation test folds the rows so every row is predicted by a fit that never saw it. The scan re-runs itself on shuffled data so you see what searching alone produces. None of these are extras — they're why the answers mean anything. What each file does touchstone.html — the main one. Four modes: Test a relation — does one quantity predict another? Fits five shapes, folds the rows, leads with the error. Try everything — 43 forms × every quantity, with a shuffle null and a bare-quantity control. My own numbers — paste two columns. Name a number — what fraction or constant is this, and can you actually tell it from its rivals? sequence_predictor.html — what comes next, and is there a rule at all? Exact rational arithmetic. Measured: 1,257 exact verdicts, 1,257 correct blind predictions, zero wrong. 800 runs of pure noise produced no verdict at all. predicted_vs_actual.html and domain_scan.html — the comparison and the wide search, standalone. What's in it 14 domains, 613 rows — metallic means, polygons, Platonic solids, phyllotaxis, the earthquake scale, the periodic table, stable isotopes, the solar system, the genetic code, amino acids, proteins, Fisher's irises, and the sequence library itself as data. Every column computed from a definition or taken from a named source. 52 integer sequences, each generated from its definition and checked against known identities, each with its OEIS number. Six languages — English, Français, Deutsch, Español, Italiano, 中文 — including every verdict and the decimal comma. What it refuses to do It won't tell you a number is meaningful because it looks tidy. 179.976% rounds to 180.0% and means nothing. It won't let a search pass as a discovery — 6,570 combinations produce a big number whether or not anything's there. It won't call 3/2 a match without telling you 2^(7/12) is 0.11% away and your number isn't that precise. It won't claim a sequence underlies a domain when the rows just happen to be countable. What it has actually found Kepler's law from the data, exponent 1.4978 against 3/2, and deterministic to 1.5 parts per million by the reverse-fit test. Gutenberg–Richter to 4 × 10⁻¹⁴%. λ − 1/λ = n exactly. The n-gon interior angle exactly. Your P1/P2 frequency equipartition holding at 0.19% and the mass-weighted version failing at 12%. That the Moon corrupts every orbital relation because its figures are Earth-relative. How much it's been tested Roughly 5,000 assertions across eight suites, 387 cross-language cases proving the JavaScript and Python agree field by field, a sweep over 94,170 domain-pair-form combinations, an independent audit re-deriving every statistic with numpy and scipy, and four headless browser suites. 58/58 manifest. That testing found real bugs in my own work — a slope invented from rounding noise, rows silently vanishing from results, a dead language function leaving a menu empty. The tool's standards applied to the tool.

View source

Similar papers

#computer vision Review Open access May 2015

A survey study on major technical barriers affecting the decision to adopt cloud services

The comparison of adopter and non-adopter sample reveals three potential adoption inhibitor, security, data privacy, and portability, which underlines the importance of the technical and security perspectives for research investigating the adoption of technology.

Nattakarn Phaphoom, Xiaofeng Wang, S. Samuel et al. · 111 citations · ⚡8
#computer vision Open access Feb 2018

Lean Internal Startups for Software Product Innovation in Large Companies: Enablers and Inhibitors

This study investigates how Lean internal startup facilitates software product innovation in large companies and identifies its enablers and inhibitors, and shows the potential of the method-in-action framework to investigate the Lean startup approach in non-startup context.

Henry Edison, Nina M. Smørsgård, Xiaofeng Wang et al. · 78 citations · ⚡6
#computer vision Book Open access Jul 2015

Understanding the affect of developers: theoretical background and guidelines for psychoempirical software engineering

This paper highlights the challenges to conduct proper affect-related studies with psychology, provides a comprehensive literature review in affect theory, and proposes guidelines for conducting psychoempirical software engineering.

D. Graziotin, Xiaofeng Wang, P. Abrahamsson · 56 citations · ⚡4
#machine learning Open access May 2017

What Influences the Speed of Prototyping? An Empirical Investigation of Twenty Software Startups

This study conducts a multiple case study on twenty European software startups and proposes a prototype-centric learning model in early stage software startups, and identifies factors that occur as barriers but also facilitators for prototyping in earlystage software startups.

Anh Nguyen-Duc, Xiaofeng Wang, P. Abrahamsson · 44 citations · ⚡5
#protein folding Open access Sep 2026

Programmable design of functional proteins from natural language

Pinal, a 16-billion-parameter foundation model that produces protein candidates from natural-language functional descriptions, supports natural language as a high-level interface for candidate generation in protein design, enabling programmable exploration with reduced reliance on manually specified structural or sequence constraints.

Fengyuan Dai, Shiyang You, Yudian Zhu et al. · 31 citations · ⚡3

Related blog posts

MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

Google DeepMind Blog Nov 25, 2025

AlphaFold: Five years of impact

Explore how AlphaFold has accelerated science and fueled a global wave of biological discovery.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.