Skip to content
#small language model Open access

Errors Not Yet Found Can Be Counted from Two Readers' Overlap, but Come Out Too Few If Both Miss the Same Ones ── If two readers find 20 and 25 errors with 10 in common, the total is estimated at 50.000000 and the errors not yet found at 15.000000 ── if the readers' blind spots are opposite, 27 are reported as 155.571429 ── [Paper 585]

Sep 2026 · Zenodo (CERN European Organization for Nuclear Research)

Abstract

One way to estimate how many errors remain in a manuscript is to have two people read it separately and use the overlap of what they found. What this paper shows is that the estimate is right only when the two readers' misses are unrelated; if they miss the same errors it comes out too few, if they miss opposite ones too many, and which way it errs is set by the sign of the covariance of their misses. No new theorem or law is claimed. Scope of this paper (scope note): No new theorem or law is claimed──the Lincoln--Petersen estimate and the Chapman correction are standard in capture--recapture. Only expected values are computed──the variation from one trial to another and confidence intervals are not treated. Findability is split into two representative groups──half at 0.7 and half at 0.1; the distribution in real manuscripts is not stated. Only two readers are considered. No proofreading procedure is designed──who should read how many times is not stated. Relation to earlier papers: Paper 334 showed that combining independent measurements adds precisions, but that perfectly correlated sensors are one──here too the assumption of independence decides the number: there correlation broke the addition of precisions, here it breaks the estimated total, and by its direction the estimate errs either way. Paper 579 showed that a correct-response rate is not a property of the question alone but a relation between question and group──here too, that findability differs from error to error is what moves the number. Paper 300 showed that sharing a root is decidable──coming out too few and coming out too many share one root, the covariance of the misses. What is added is confirming by expected values that the estimate is exactly right when findability is uniform, setting out that shared misses report 45 unfound errors as 9.000000 and opposite misses report 27 as 155.571429, placing the error in a covariance formula checked against the direct calculation, and placing the separator on the sign of the covariance. First, the overlap gives the total──if A finds 20, B finds 25 and both find 10, the total is estimated at 50.000000 and the errors not yet found at 15.000000; the small-sample correction gives 48.636364 (Section 2). Second, if findability is uniform, the estimate is exactly right──with 100 true errors the expected estimate is 100.000000, and the estimate of those not found, 30.000000, matches the true 30 (Section 3). Third, and this is the core. If both readers miss the same errors, the estimate comes out too few──when half the errors are found by each reader with probability 0.7 and half with 0.1, the overlap swells to 25.0000, the total is estimated at 64.000000, and the 45 errors not yet found are reported as 9.000000 (Section 4). Fourth, if their blind spots are opposite, the estimate comes out too many──the overlap shrinks to 7.0000, the total is estimated at 228.571429, and 27 errors not yet found are reported as 155.571429 (Section 5). Fifth, the error of the estimate has one formula──the ratio of estimate to truth is 1/(1+cov/( p_1 p_2)), matching the direct calculation to 10^-12 (Section 6). Sixth, the separator is whether the two readers' misses are related, and in which direction──a positive covariance gives too few, a negative one too many, and only zero is right (Section 6). One way to estimate how many errors remain in a manuscript is to have two people read it separately and use the overlap of what they found. The number of the unseen comes from the overlap of the seen──if A finds 20, B 25 and both 10, the total is estimated at 50.000000 and 15.000000 remain, or 48.636364 with the small-sample correction; with uniform findability the expected estimate hits the true 100 exactly. But if both readers miss the same errors, the estimate comes out too few──with half the errors easy for both and half hard for both, the overlap swells to 25.0000, the total is estimated at 64.000000, and the 45 not yet found are reported as 9.000000: a large overlap is misread as "almost all found". With opposite blind spots the overlap shrinks to 7.0000 and 27 unfound errors are reported as 155.571429. The separator is the sign of the covariance of the two readers' misses──the ratio of estimate to truth is 1/(1+cov/( p_1 p_2)): too few for a positive covariance, too many for a negative one, right only at zero. Whether the estimate is right cannot be told from the counts; it rests on an unseen property, whether the two are weak on the same errors. Under Paper 300 the two directions of error share one root. Placed among the earlier papers──as in Paper 334 independence decides the number, and here its direction makes the estimate err either way. As in Paper 579, non-uniformity moves the number. To be honest──only expected values were computed, with no variation or confidence intervals; findability was split into two representative groups and only two readers considered. And no proofreading procedure is designed: what is shown is on what assumption the estimate is right, and which way it errs. On the making of this work: The ideas and content of this work stem from the author's own considerations. Assistance from an AI (a large language model) was used for structuring, English translation, and checking the algebra. Any remaining errors or misinterpretations are solely the author's. Feedback and corrections are sincerely appreciated. Keywords: capture-recapture, Lincoln-Petersen estimator, proofreading, number of errors, independence assumption. ----- 原稿を二人が別々に読み、見つけた誤りの重なりから、まだ見つかっていない誤りの数を推定する方法がある。本稿が示すのは、その推定は二人の見落としが互いに関係しないときだけ正しく、同じ所で見落とせば少なく、逆の所で見落とせば多く出て、どちらに外れるかは見落としの共分散の符号で決まることである。新しい定理も法則も主張しない。 本稿の射程(射程注記):新しい定理も法則も主張しない──リンカーン=ピーターセンの推定も、チャップマンの補正も、捕獲再捕獲の標準である。期待値だけを計算した──一回ごとのばらつきや信頼区間は扱わない。誤りの見つけやすさは二つの組に分けた代表値である──半分は 0.7、半分は 0.1 という分け方で、実際の原稿での分布は述べない。読み手は二人に限った。校正の手順を設計しない──誰に何回読ませるべきかは述べない。既刊との関係:論文334 は、独立な測りを重ねると精度が足し算になるが、完全に相関していれば三台は一台だと示した──本稿でも独立の仮定が数を決める。あちらは相関で精度の足し算が崩れ、こちらは相関で推定の総数が崩れる。しかも相関の向きで、少なくも多くも外れる。論文579 は、正答率が問題だけの性質ではなく、問題と集団の関係であると示した──本稿でも、誤りごとに見つけやすさが違うことが数を動かす。論文300 は同根か別根かが判定できると示した──少なく出ることと多く出ることは、見落としの共分散という一つの根から出る同根である。加えたのは、見つけやすさが一様なら推定がちょうど当たることを期待値で確かめたこと、同じ所で見落とすと見つかっていない 45 を 9.000000 と答え、逆の所で見落とすと 27 を 155.571429 と答えることを並べたこと、外れ方を共分散の式に置き、直接の計算と突き合わせたこと、分離子を共分散の符号に置いたことである。 第一に、重なりから総数が出る──A が 20、B が 25、両方が 10 見つけたなら、推定の総数は 50.000000、見つかっていないのは 15.000000 である。小さな標本の補正を入れると 48.636364 になる(第2節)。 第二に、見つけやすさが一様なら、推定はちょうど当たる──本当の総数 100 に対し、期待値で推定の総数は 100.000000、見つかっていない推定 30.000000 も本当の 30 に一致する(第3節)。 第三に、これが本稿の芯である。二人が同じ所で見落とすと、推定は少なく出る──半分の誤りは二人とも 0.7 で、半分は 0.1 で見つけるとき、重なりは 25.0000 に膨らみ、推定の総数は 64.000000、見つかっていない 45 を 9.000000 と答える(第4節)。 第四に、見落とす所が二人で逆なら、推定は多く出る──重なりは 7.0000 に縮み、推定の総数は 228.571429、見つかっていない 27 を 155.571429 と答える(第5節)。 第五に、外れ方は一つの式で書ける──推定と本当の比は 1/(1+cov/( p_1 p_2)) で、直接の計算と 10^-12 で一致する(第6節)。 第六に、分離子は、二人の見落としが互いに関係するかどうか、とその向きである──共分散が正なら少なく、負なら多く、0 のときだけ当たる(第6節)。 原稿を二人が別々に読み、見つけた誤りの重なりから、まだ見つかっていない誤りの数を推定する方法がある。見えないものの数が、見えたものの重なりから出る──A が 20、B が 25、両方が 10 見つけたなら、推定の総数は 50.000000、見つかっていないのは 15.000000 で、小さな標本の補正を入れると 48.636364 になる。見つけやすさが一様なら、期待値で推定はちょうど本当の 100 に当たる。ところが二人が同じ所で見落とすと、推定は少なく出る──半分の誤りは二人とも見つけやすく、半分は二人とも見つけにくいとき、重なりは 25.0000 に膨らみ、推定の総数は 64.000000、見つかっていない 45 を 9.000000 と答える。重なりが大きいことを「もうほとんど見つけた」と読み違える。見落とす所が二人で逆なら、重なりは 7.0000 に縮み、見つかっていない 27 を 155.571429 と答える。分離子は、二人の見落としの共分散の符号である──推定と本当の比は 1/(1+cov/( p_1 p_2)) で、共分散が正なら少なく、負なら多く、0 のときだけ当たる。推定が正しいかどうかは見つけた数からは分からず、二人が同じ種類の誤りに弱いかという見えない性質で決まる。論文300 の基準で、少なく出ることと多く出ることは同根である。既刊との位置──論文334 と同じく独立の仮定が数を決め、こちらでは相関の向きで少なくも多くも外れる。論文579 と同じく、一様でないことが数を動かす。正直に言えば──期待値だけを計算し、ばらつきも信頼区間も扱っていない。見つけやすさは二つの組に分けた代表値で、読み手は二人に限った。そして校正の手順は設計しない──示したのは、重なりからの推定がどの前提で当たり、どちらに外れるかである。 作成にあたって:本稿の着想と内容は、著者自身の考察に基づくものです。文章の構成整理や英訳、数式の確認には AI(大規模言語モデル)の助力を得ました。最終的な内容の解釈や誤りがあれば、それらはすべて著者の責に帰します。お気づきの点があれば、ご教示いただければ幸いです。 キーワード:捕獲再捕獲法、リンカーン=ピーターセン推定、校正の見落とし、誤りの数、独立の仮定。

View source

Similar papers

#small language model Dataset Open access Oct 2026

Socratic guiding questions in synthetic arithmetic data: matched LoRA runs (revision v2)

Supporting data, adapters, predictions and code for the article *Low-Cost LoRA Fine-Tuning of Small Language Models for Multi-Step Arithmetic Reasoning* by Jake O'Grady, Asena Isik Gürhan, Chee Fong Ting and Effirul Ramlan (University of Galway). We generated 20,000 GSM8K-derived arithmetic problems with step-by-step s...

O'Grady, Jake, Gürhan, Asena Isik, Chee, Fong Ting et al. · 465 citations
#computer vision Open access Jun 2016

Software Development in Startup Companies: The Greenfield Startup Model

The results are packaged in the Greenfield Startup Model (GSM), which explains the priority of startups to release the product as quickly as possible, and the need to shorten time-to-market, by speeding up the development through low-precision engineering activities.

Carmine Giardino, Nicolò Paternoster, M. Unterkalmsteiner et al. · 178 citations · ⚡14
#computer vision Open access Oct 2016

Software Startups - A Research Agenda

Software startup companies develop innovative, software-intensive products within limited timeframes and with few resources, searching for sustainable and scalable business models.

M. Unterkalmsteiner, P. Abrahamsson, Xiaofeng Wang et al. · 157 citations · ⚡17
#machine learning Review Open access Oct 2016

“Failures” to be celebrated: an analysis of major pivots of software startups

This study conducts a case survey study based on the secondary data of the major pivots happened in 49 software startups, and demonstrates that customer need pivot is the most common among all pivot types.

Sohaib Shahid Bajwa, Xiaofeng Wang, Anh Nguyen-Duc et al. · 127 citations · ⚡15
#computer vision Review Open access May 2015

A survey study on major technical barriers affecting the decision to adopt cloud services

The comparison of adopter and non-adopter sample reveals three potential adoption inhibitor, security, data privacy, and portability, which underlines the importance of the technical and security perspectives for research investigating the adoption of technology.

Nattakarn Phaphoom, Xiaofeng Wang, S. Samuel et al. · 111 citations · ⚡8
#computer vision Open access Feb 2018

Lean Internal Startups for Software Product Innovation in Large Companies: Enablers and Inhibitors

This study investigates how Lean internal startup facilitates software product innovation in large companies and identifies its enablers and inhibitors, and shows the potential of the method-in-action framework to investigate the Lean startup approach in non-startup context.

Henry Edison, Nina M. Smørsgård, Xiaofeng Wang et al. · 78 citations · ⚡6

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.