Skip to content
#protein folding Open access

Two Millennium Problems for Biology, reduced to arithmetic: what the acceptance criteria for Problems 3 and 10 already force

Sep 2026 · Zenodo (CERN European Organization for Nuclear Research) · 4 references

Abstract

Start with the six-page PDF. On 19 September 2026 Edison Scientific and FutureHouse published twelve Millennium Problems for Biology with explicit acceptance criteria. This note takes two of them, Problem 3 (the reverse translatase) and Problem 10 (the protein amplification chain reaction), and proves what the criteria themselves force. 1. Problem 3: the degeneracy objection is aimed at the wrong direction. Translation of a coding segment is a surjection onto peptides, so it has a right inverse: for every peptide a coding sequence exists, and one fixed choice of codon per amino acid produces it. That fixed choice is exactly what Problem 3 calls a preregistered codon convention. Separately, the historical codons are not recoverable from the peptide: the fibre has size the product of the codon multiplicities, the optimal worst-case probability of naming the original is its reciprocal, and exact recovery needs a side message of ceil(log2 of the fibre size) bits. For a 50-residue worst case that is 650 = 8.08 x 1038 candidate genes and exactly 130 bits. Both statements are proved. Together they say that the standard objection to reverse translation, that the code is degenerate and therefore not invertible, is true about reading history and irrelevant to Problem 3, which asks the enzyme to write under a fixed convention and then asks for the peptide back, not the original gene. What is left is chemistry. 2. Problem 10: one threshold, a length ceiling, and a survival probability. For any amplifier passing through a single intermediate carrier, with constant per-molecule rates and inert errors, the two-state system grows exponentially if and only if ab > muP muM, with the growth rate given in closed form. Nothing in that computation asks what the carrier is made of, so Problem 10's prohibition on nucleic-acid intermediates changes no formula. If reading destroys the parent, the destruction rate enters the parent's own loss term and a finite reading rate exists if and only if aq > theta m, with an exact critical rate; under independent per-position errors this converts a fidelity into a hard ceiling on peptide length. From a finite seed the answer is a probability rather than a certainty: the two-type branching process has exact survival probabilities, given in closed form. 3. Two numbers Problem 10 fixes about itself. At the 50-mer the criteria specify, and with erroneous copies inert, break-even requires R0 > h-50 daughter attempts per parent per cycle: about 194 at 90 percent per-residue fidelity, about 1.65 at 99 percent. And because errors compound along a lineage, the criteria's own pair of numbers, 1000-fold amplification and 90 percent per-residue accuracy, already imply a per-cycle fidelity floor: 98.83 percent for a doubling chemistry, 97.76 percent at a gain of four, 96.41 percent at a gain of ten. These are necessary conditions on the whole product pool, derived from the fact that at most bj copies of lineage depth at most j can ever exist. Claim boundary. No enzyme is constructed, no amplification is demonstrated, and neither Problem 3 nor Problem 10 is solved or partially solved in the sense of its acceptance criteria. Sections 2 to 5 analyse a model whose rate constants are assumed, not measured. The mathematics is elementary and established; Section 6 records the prior art, including Craig 1981, the Martin patent, Shi et al. 2026 and Zheng et al. 2026, and claims no priority. Every number in the note is recomputed from scratch by the deposited script check_numbers.py, which also checks Theorem 3 against 20000 random rate sets and Theorem 5 against direct fixed-point iteration.

View source

Similar papers

#computer vision Review Open access May 2015

A survey study on major technical barriers affecting the decision to adopt cloud services

The comparison of adopter and non-adopter sample reveals three potential adoption inhibitor, security, data privacy, and portability, which underlines the importance of the technical and security perspectives for research investigating the adoption of technology.

Nattakarn Phaphoom, Xiaofeng Wang, S. Samuel et al. · 111 citations · ⚡8
#computer vision Open access Feb 2018

Lean Internal Startups for Software Product Innovation in Large Companies: Enablers and Inhibitors

This study investigates how Lean internal startup facilitates software product innovation in large companies and identifies its enablers and inhibitors, and shows the potential of the method-in-action framework to investigate the Lean startup approach in non-startup context.

Henry Edison, Nina M. Smørsgård, Xiaofeng Wang et al. · 78 citations · ⚡6
#computer vision Book Open access Jul 2015

Understanding the affect of developers: theoretical background and guidelines for psychoempirical software engineering

This paper highlights the challenges to conduct proper affect-related studies with psychology, provides a comprehensive literature review in affect theory, and proposes guidelines for conducting psychoempirical software engineering.

D. Graziotin, Xiaofeng Wang, P. Abrahamsson · 56 citations · ⚡4
#machine learning Open access May 2017

What Influences the Speed of Prototyping? An Empirical Investigation of Twenty Software Startups

This study conducts a multiple case study on twenty European software startups and proposes a prototype-centric learning model in early stage software startups, and identifies factors that occur as barriers but also facilitators for prototyping in earlystage software startups.

Anh Nguyen-Duc, Xiaofeng Wang, P. Abrahamsson · 44 citations · ⚡5
#protein folding Open access Sep 2026

Programmable design of functional proteins from natural language

Pinal, a 16-billion-parameter foundation model that produces protein candidates from natural-language functional descriptions, supports natural language as a high-level interface for candidate generation in protein design, enabling programmable exploration with reduced reliance on manually specified structural or seque...

Fengyuan Dai, Shiyang You, Yudian Zhu et al. · 31 citations · ⚡3

Related blog posts

Google DeepMind Blog Sep 30, 2026

Introducing SynthID Bio

Proof of concept for watermarking AI-generated proteins while preserving biological function.

MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.