HELD (_) -- Reproducibility deposit
Abstract
A guaranteed-coverage confidence interval for the two-sample standardized effect William J. Dwyer, MD, MPH, FAAP — Department of Mathematics and Statistics, University of Massachusetts Lowell. ORCID 0009-0004-0855-7222. Concept DOI (always resolves to the latest version): minted on first publication. What this is The reproducibility deposit for the m01te methods paper: a guaranteed-coverage confidence interval for the two-sample standardized effect (Cohen's d) at the skewed, unequal-variance, small-n corner where the textbook interval silently under-covers. The noncentral-t inversion assumes normal data and equal variances; at a lognormal, four-to-one variance-ratio, n = 10 design its realized coverage falls to 0.81 against a nominal 0.95, and a naive percentile bootstrap of dfalls further, to 0.78 — a joint failure of the mean-difference reference and the variance estimate that standardizes it, which resampling does not repair. The paper gives a three-tier recommendation, mirroring the companion two-sample test and the one-way effect-size paper: Classical — the noncentral-t / normal-approximation interval, the everyday default, liberal at the corner. Calibrated middle tier — the guaranteed two-sample test T_BB inverted for the mean difference at the full level, divided by the plug-in pooled scale. Closed-form and deterministic (no resampling), with near-nominal worst-case coverage 0.93 at about 1.6× the classical width. It keeps the mean-difference deflation that repairs the actual under-coverage while treating the scale at its point estimate; the over-covering numerator and the under-covering plugged-in scale roughly cancel to near nominal. Guaranteed floor — a Bonferroni combination of the T_BB-inverted mean-difference interval with a distribution-free bootstrap scale interval, carrying a proved finite-sample coverage floor (worst-case 0.97) at about 4× the classical width. What the deposit contains Manuscript (author + anonymized markdown; built .docx/.pdf, including a cross-reference–hyperlinked variant) and the derivations (D1–D6): the estimand and its d_av scale; the T_BB-inverted mean-difference interval; the exact Bonferroni coverage floor of the ratio interval; why the classical standard error under-covers off its normal/equal-variance premise; the deterministic-simulation confirmation; and the calibrated middle tier with its compensation argument. Reproducibility runner — rerun/rc_m01te_coverage.py computes, for each design cell across the parent-distribution × sample-size × variance-ratio × effect grid, the realized coverage and mean width of all four intervals (classical, percentile-bootstrap, calibrated middle, guaranteed floor). Every number regenerates from this deterministically-seeded script (seed 20260826); its locked output CSV is deposited. Figure — figures/m01te_coverage.png (built by make_m01te_figure.py): the four coverage curves cell by cell across the grid, the classical and bootstrap curves sliding below nominal at the corner, the calibrated curve tracking near it, and the guaranteed curve holding above it. All evaluation is simulation-based. Code is released under the MIT License; text and figures under CC BY 4.0. How to cite Please cite this deposit if you use the package or the method. Citing the concept DOI references the work in general and always resolves to the latest version; cite a specific version DOI to point at an exact snapshot. Dwyer, W. J. (2026). A guaranteed-coverage confidence interval for the two-sample standardized effect: reproducibility deposit (Version 1.0.0) [Software]. Zenodo. https://doi.org/⟨concept DOI⟩