Skip to content

HELD (_) -- Reproducibility deposit

Aug 2026 · Zenodo (CERN European Organization for Nuclear Research)

Abstract

A guaranteed-coverage confidence interval for the two-sample standardized effect William J. Dwyer, MD, MPH, FAAP — Department of Mathematics and Statistics, University of Massachusetts Lowell. ORCID 0009-0004-0855-7222. Concept DOI (always resolves to the latest version): minted on first publication. What this is The reproducibility deposit for the m01te methods paper: a guaranteed-coverage confidence interval for the two-sample standardized effect (Cohen's d) at the skewed, unequal-variance, small-n corner where the textbook interval silently under-covers. The noncentral-t inversion assumes normal data and equal variances; at a lognormal, four-to-one variance-ratio, n = 10 design its realized coverage falls to 0.81 against a nominal 0.95, and a naive percentile bootstrap of dfalls further, to 0.78 — a joint failure of the mean-difference reference and the variance estimate that standardizes it, which resampling does not repair. The paper gives a three-tier recommendation, mirroring the companion two-sample test and the one-way effect-size paper: Classical — the noncentral-t / normal-approximation interval, the everyday default, liberal at the corner. Calibrated middle tier — the guaranteed two-sample test T_BB inverted for the mean difference at the full level, divided by the plug-in pooled scale. Closed-form and deterministic (no resampling), with near-nominal worst-case coverage 0.93 at about 1.6× the classical width. It keeps the mean-difference deflation that repairs the actual under-coverage while treating the scale at its point estimate; the over-covering numerator and the under-covering plugged-in scale roughly cancel to near nominal. Guaranteed floor — a Bonferroni combination of the T_BB-inverted mean-difference interval with a distribution-free bootstrap scale interval, carrying a proved finite-sample coverage floor (worst-case 0.97) at about 4× the classical width. What the deposit contains Manuscript (author + anonymized markdown; built .docx/.pdf, including a cross-reference–hyperlinked variant) and the derivations (D1–D6): the estimand and its d_av scale; the T_BB-inverted mean-difference interval; the exact Bonferroni coverage floor of the ratio interval; why the classical standard error under-covers off its normal/equal-variance premise; the deterministic-simulation confirmation; and the calibrated middle tier with its compensation argument. Reproducibility runner — rerun/rc_m01te_coverage.py computes, for each design cell across the parent-distribution × sample-size × variance-ratio × effect grid, the realized coverage and mean width of all four intervals (classical, percentile-bootstrap, calibrated middle, guaranteed floor). Every number regenerates from this deterministically-seeded script (seed 20260826); its locked output CSV is deposited. Figure — figures/m01te_coverage.png (built by make_m01te_figure.py): the four coverage curves cell by cell across the grid, the classical and bootstrap curves sliding below nominal at the corner, the calibrated curve tracking near it, and the guaranteed curve holding above it. All evaluation is simulation-based. Code is released under the MIT License; text and figures under CC BY 4.0. How to cite Please cite this deposit if you use the package or the method. Citing the concept DOI references the work in general and always resolves to the latest version; cite a specific version DOI to point at an exact snapshot. Dwyer, W. J. (2026). A guaranteed-coverage confidence interval for the two-sample standardized effect: reproducibility deposit (Version 1.0.0) [Software]. Zenodo. https://doi.org/⟨concept DOI⟩

View source

Similar papers

#computer vision Review Sep 2017

Agile Software Development Methods: Review and Analysis

This publication proposes a definition and a classification of agile software development approaches and analyses ten software development methods that can be characterized as being "agile" against the defined criterion.

P. Abrahamsson, O. Salo, Jussi Ronkainen et al. · 727 citations · ⚡54
#computer vision Jun 2008

The impact of agile practices on communication in software development

The study shows that agile practices improve both informal and formal communication, but indicates that, in larger development situations involving multiple external stakeholders, a mismatch of adequate communication mechanisms can sometimes even hinder the communication.

M. Pikkarainen, Jukka Haikara, O. Salo et al. · 401 citations · ⚡48
#machine learning Review Open access Oct 2014

Software development in startup companies: A systematic mapping study

The results indicate that software engineering work practices are chosen opportunistically, adapted and configured to provide value under the constrains imposed by the startup context.

Nicolò Paternoster, Carmine Giardino, M. Unterkalmsteiner et al. · 394 citations · ⚡54

Related blog posts

GPT-Lab Sep 17, 2026

Beyond Prompt Engineering: The Role of Tacit Knowledge in Software Engineering

AI is making software generation faster, but speed does not remove the need for expertise. As more work is delegated to AI, tacit knowledge may become one of the most important human advantages in software engineering. The post Beyond Prompt Engineering: The Role of Tacit Knowledge in Software Engineering appeared first on GPT-Lab.

Microsoft Research Blog Jul 30, 2026

Echoverse: Deep, evolving environments for computer-use agents

Computer-use AI agents struggle with multi-step workflows like email and customer support. Echoverse trains agents in realistic environments rather than simply providing more training tasks, helping them improve as the tasks, tests, and environments evolve. The post Echoverse: Deep, evolving environments for computer-use agents appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.