Skip to content

A guaranteed-level repair for Welch's two-sample t, which silently over-rejects under skew, heteroscedasticity, and small samples (m01t) — reproducibility deposit

Aug 2026 · Zenodo (CERN European Organization for Nuclear Research)

Abstract

A guaranteed-level repair for Welch's two-sample t, which silently over-rejects under skew, heteroscedasticity, and small samples William J. Dwyer, MD, MPH, FAAP — Department of Mathematics and Statistics, University of Massachusetts Lowell. ORCID 0009-0004-0855-7222. Concept DOI (always resolves to the latest version): 10.5281/zenodo.22036361. Published v1.4.3:10.5281/zenodo.22104110. What this is The reproducibility deposit for the m01t paper — the two-sample companion to the guaranteed-level ANOVA work (m01A/m01x). Welch's two-sample t is the field default for comparing two means under unequal variances, but at the corner where the data are skewed and heteroscedastic and the samples are small it runs liberal: it rejects the null more often than its nominal level allows. A scan of 515 real public-data comparisons finds that about one in three borderline-significant Welch results do not survive a level-guaranteed test — a concrete, real-data measure of the over-rejection. m01t specializes the companion procedure's Berger-Boos deflation to two groups: it deflates the known-variance quadratic by a closed-form radius built from each group's kurtosis-widened variance instability and refers the result to a chi-square. The test holds worst-case size at or below nominal exactly where Welch is liberal; its raw conservatism is the honest cost of the guarantee, and size-adjusted it matches Welch to within about 0.02 in power. The deposit ships the browser demonstrator honest_ttest.html, which computes Welch beside the guaranteed T_BB on your own data and prints the honest receipt — a deflation mechanism (the guaranteed rejection region is a strict subset of Welch's, so a significant call can only ever be withdrawn, never manufactured; no reverse flip is possible). What the deposit contains Manuscript (author + anonymized; built .docx/.pdf), the novelty / prior-art companion, and the derivationscompanion (the two-group Berger-Boos radius, the kurtosis-widening bound, the size proof, and the one-sided variant). Reproducibility apparatus — the simulation runners (the shared-grid scoreboard, the Fleishman skew×kurtosis decomposition grid, the competitor and gate-ablation runners) and the locked result CSVs. Every reported number regenerates from these deterministically-seeded scripts. Figures — the headline, the Welch over-rejection heat map, the validity-versus-power frontier, the method × stress-regime worst-size heat map, the per-method failure map, the kurtosis-axis curve, and the orthogonal-axis decomposition of Welch's realized size. Interactive demonstrator honest_ttest.html — reproduces the deposited Python exactly and carries the house flip-interpretation standard (deflation marker, "how to read a flip" beat, real-data incidence panel, and an assertNoReverseload guard). Deep-dive record (transfer/power-comparison, the 12-cell flip taxonomy with citations, the declarations and reference-block fixes, the decomposition study) and the submission apparatus. All evaluation is simulation-based. Code is released under the MIT License; text, figures, and data under CC BY 4.0. How to cite Please cite this deposit if you use the package or the method. Citing the concept DOI references the work in general and always resolves to the latest version; cite a specific version DOI to point at an exact snapshot. Dwyer, W. J. (2026). A guaranteed-level repair for Welch's two-sample t — reproducibility deposit[Software]. Zenodo. https://doi.org/10.5281/zenodo.22036361 BibTeX: bibtex @software{dwyer_m01t_2026, author = {Dwyer, William J.}, title = {A guaranteed-level repair for Welch's two-sample t --- reproducibility deposit}, year = {2026}, publisher = {Zenodo}, doi = {10.5281/zenodo.22036361}, url = {https://doi.org/10.5281/zenodo.22036361}, orcid = {0009-0004-0855-7222} } The DOI above is the concept DOI (resolves to the latest version); to cite a specific release use that version's DOI in place of it (e.g. 10.5281/zenodo.22104110 for v1.4.3). When the accompanying journal article appears, please cite it as the primary reference for the method and this deposit as the reproducibility archive. Version history v1.4.4 — ✅ 10.5281/zenodo.22187004 (2026-08-31) staged (pending upload): rendering-only refresh — the manuscript and Derivations docx/PDF rebuilt through the current math-typography builder so nested-paren radicals draw as true Office-math radicals; the deposit now also carries the current shared house tooling (the bundled mseries_deposit.py includes the require_all deposit guard). No number, figure, table, or claim changed. New version on concept 10.5281/zenodo.22036361 (m01t_reproducibility_v1.4.4.zip, md5 6e4e09e7dd1a5a747fa1493db94c922d, 3,561,269 B, 73 files). v1.4.3 — ✅ 10.5281/zenodo.22104110 (2026-08-24): arXiv source refreshed; declarations split into separate paragraphs; reference-block one-per-line fix. v1.4.2 — ✅ 10.5281/zenodo.22103410 · v1.4.1 — ✅ 10.5281/zenodo.22102558 · v1.4.0 (2026-08-23): new Figure 8, the orthogonal skew/kurtosis decomposition of Welch's realized size (Fleishman power method; kurtosis alone does not inflate the two-sided size, T_BB holds throughout). v1.3.0 — ✅ 10.5281/zenodo.22054950 (2026-08-22): deposit-scale shared-grid contest fold (40,000 reps, B=699); Table 6 becomes a five-method property scorecard; four contest figures added. v1.1.0 — ✅ 10.5281/zenodo.22036735 (2026-08-21): reviewer- hardening leads folded (held-out-family validity, Monte-Carlo CI on the worst-case size, studentized bootstrap-t and permutation benchmarks, the one-sided T_BB). v1.0.0 — first deposit (concept 10.5281/zenodo.22036361): the guaranteed-level two-group repair, the 515-comparison public-data flip scan, and the honest_ttest.html demonstrator. Provenance: every number traces to a named, deterministically-seeded runner; the demonstrator reproduces the deposited Python (parity self-checked on load). Related records: the one-way ANOVA sibling m01A 10.5281/zenodo.21908169; the moderated-Welch record m01x; the T_root methodology 10.5281/zenodo.21522471.

View source

Similar papers

#computer vision Review Sep 2017

Agile Software Development Methods: Review and Analysis

This publication proposes a definition and a classification of agile software development approaches and analyses ten software development methods that can be characterized as being "agile" against the defined criterion.

P. Abrahamsson, O. Salo, Jussi Ronkainen et al. · 727 citations · ⚡54
#computer vision Jun 2008

The impact of agile practices on communication in software development

The study shows that agile practices improve both informal and formal communication, but indicates that, in larger development situations involving multiple external stakeholders, a mismatch of adequate communication mechanisms can sometimes even hinder the communication.

M. Pikkarainen, Jukka Haikara, O. Salo et al. · 401 citations · ⚡48
#machine learning Review Open access Oct 2014

Software development in startup companies: A systematic mapping study

The results indicate that software engineering work practices are chosen opportunistically, adapted and configured to provide value under the constrains imposed by the startup context.

Nicolò Paternoster, Carmine Giardino, M. Unterkalmsteiner et al. · 394 citations · ⚡54

Related blog posts

GPT-Lab Sep 17, 2026

Beyond Prompt Engineering: The Role of Tacit Knowledge in Software Engineering

AI is making software generation faster, but speed does not remove the need for expertise. As more work is delegated to AI, tacit knowledge may become one of the most important human advantages in software engineering. The post Beyond Prompt Engineering: The Role of Tacit Knowledge in Software Engineering appeared first on GPT-Lab.

Microsoft Research Blog Jul 30, 2026

Echoverse: Deep, evolving environments for computer-use agents

Computer-use AI agents struggle with multi-step workflows like email and customer support. Echoverse trains agents in realistic environments rather than simply providing more training tasks, helping them improve as the tasks, tests, and environments evolve. The post Echoverse: Deep, evolving environments for computer-use agents appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.