Conformal Effort Intervals: Distribution-Free Prediction Intervals for Software Effort and Their Empirical Reliability under Deployment Shift
Abstract
Replication package for Conformal Effort Intervals: Distribution-Free Prediction Intervals for Software Effort and Their Empirical Reliability under Deployment Shift. Conformal prediction turns any effort estimator into one that outputs an interval with a coverage guarantee that is distribution-free and holds for any base learner. Computing the nonconformity score in log space makes the interval multiplicative, ŷ×[1/k, k], which matches the right-skewed geometry of effort: a task predicted at 10 hours with k = 2.5 gives [4, 25] hours. The guarantee is marginal and conditional on exchangeability, which is what this study sets out to test. What the study found Across eight datasets, from four classical benchmarks to 59,934 tasks carrying a planner estimate: Under random splits, realized coverage is broadly consistent with nominal across four base learners and every level the calibration sets support Holding out an entire organization, a model without a human estimate covers 72.14% against a 90% target, with a worst fold of 62.07%, while a conformalized planner estimate averages 90.06% That failure is a calibration-transfer failure, not an irreducible limit: calibrating on 1,000 tasks from the target organization recovers the model arm to 90.22%, at roughly triple the interval width One column distributed with the China benchmark raises the share of estimates within 25% of actual from 24.24% to 96.04%, which an implausibly narrow interval exposes Against a historical empirical quantile, the finite-sample correction is the whole contribution, and it matters only when the calibration set is small: 90.34% against 85.51% at 20 calibration points, indistinguishable by 500 Contents src/ — the conformal module: split conformal, cross-conformal, CV+ and CQR, each documented with the guarantee it actually carries notebooks/ — six executable notebooks, including the definitive CESAW run results/ — every stored table and figure, with the script that produced each results/controlled/ — five controlled comparisons that isolate a confound in the main experiments: Mondrian calibration with the model held fixed, a common calibration sample across predictor arms, identifier coding, target-organization calibration, and the historical-quantile baseline results/superseded/ — earlier outputs kept for provenance; none backs a reported number data/raw/ — the four classical benchmarks, with provenance and leakage notes Verification Three scripts run here, and VERIFICATION.md records what each covers and what it does not. results/classical/check_reproduction.py re-runs the classical pipeline and compares its output against the stored tables; all six matched exactly at 20 repeats. check_sync.py fails if a notebook has fallen behind the committed modules. check_docs.py fails if a superseded figure reappears in the documentation. The dataset source is pinned to an upstream commit recorded in UPSTREAM-DATA-COMMIT.txt and checked out by the notebooks. Two analysis defects found after results had been produced are documented rather than removed. Requirements Python 3.10+, with requirements-lock.txt giving the exact versions used. The large datasets are cloned at runtime; the SEERA workbook must be downloaded separately from Zenodo. The full CESAW run takes about two hours.