Touchstone
Abstract
Touchstone alpha v1.0 What it is Touchstone is a reproducible statistical and algebraic screening tool for distinguishing supported mathematical relationships from search-induced patterns in sequences and numerical datasets. A methodological analysis tool for testing whether apparent mathematical structure in data is genuinely supported, or whether it is an artefact of searching many possible transformations and relationships. At its core, it does two related jobs. First, on the sequence side, it tests whether a proposed recurrence or rule has actually earned the right to be called a pattern. Rather than rewarding a formula simply because it fits the terms used to discover it, Touchstone requires out-of-sample confirmations: a rule must correctly predict further terms that were not needed to fit it. This is an explicit safeguard against overfitting and “guessing” structure from short sequences. Second, on the table-scan side, Touchstone searches numerical datasets for unusually strong relationships between columns and transformed versions of those columns. It can examine raw quantities, reciprocals and other predefined forms, but it does not treat the strongest observed correlation as meaningful by itself. Instead it passes every candidate through a sequence of null tests designed to measure how easily an equally impressive result could arise by chance after searching the same field. The scan uses three main safeguards. A fixed-hit permutation test asks whether the winning relationship remains unusual when the data are shuffled. A more conservative edge null asks how extreme the winner is relative to the strongest result found anywhere in the entire shuffled search, thereby accounting for the fact that the software looked across many column pairs and forms. Finally, a raw-quantity control distinguishes cases where a transformed equation adds something from cases where the underlying raw variables already explain the signal. The final verdict is deliberately simple. Touchstone can report an equation when the transformed relationship survives the statistical gates and improves materially on the raw relationship; raw when the underlying quantities carry the signal without justification for elevating a transformed formula; or no claim when the evidence is insufficient. The purpose is not to manufacture equations, but to make it comparatively difficult for an attractive accidental pattern to be promoted into one. A significant part of the project is therefore calibration rather than discovery. The current version has been run end-to-end on 500 synthetic pure-noise domains, each with eight columns and sixty rows. Using exactly the same search and verdict procedure as for real data, it produced equation verdicts in 4.8% of noise domains, raw verdicts in 0.4%, and no claim in the remaining 94.8%. That experiment provides an empirical type-I calibration for the complete procedure rather than assuming that two nominal 5% gates behave independently. Touchstone also measures the effective search width rather than merely counting every algebraic form as a separate opportunity. Many transformations are rank-identical on a given dataset, so the software records how many genuinely distinct ranked forms were searched. This makes the multiple-search correction more interpretable and avoids exaggerating the size of the search space. The project is intentionally conservative. Tied columns are permitted by the permutation test but are flagged as an interpretive limitation; guarded inversions are inherited from the mathematical definitions rather than introduced opportunistically for scanning; very small permutation probabilities are reported at the correct finite-simulation floor; and numerical ratio names such as 3/2 are accompanied by a measured false-naming rate. In practical terms, Touchstone is best thought of as a pattern-claim filter. You give it a short sequence or a numerical table, and it asks: Is there enough evidence to say that this is an actual rule or equation, rather than a pattern that looks convincing because we searched until something appeared? Its contribution is not the general idea of over-determination, which is well established in recurrence guessing and statistical model checking, but the combination of a clear verdict vocabulary, search-aware null testing, raw controls, and empirical calibration of how often the complete procedure makes a claim on pure noise.