Randomized Fair Policy Selection from Partially Identifying Data: Code and Numerical Data
Abstract
This version 4 reproducibility package accompanies Randomized Fair Policy Selection from Partially Identifying Data by Ayush Ojha. The paper studies selecting one policy from a fixed menu, or abstaining, under a uniform bound on the probability of an unfair output. It characterizes attainable return under shared observation laws and unsafe-pattern closure constraints, with a confidence-region construction under the stated assumptions. The package contains a deterministic integration study over 225 conditions and five rules, with six finite unsafe-pattern optimization examples. It also contains a separate retained-action-identifier study over 70 fixed synthetic conditions with 20,000 paired samples per condition: 1.4 million paired logged samples and 4.2 million method outputs. The latter compares coarsened marginal selection, a disclosed Seldonian sample-split/Clopper-Pearson instantiation, and simultaneous per-policy certification. The procedures differ in the observation content they use. The results distinguish useful safe return, unsafe return, and total return; they do not establish a universal method ranking or real-world fairness. Included materials are research scripts, pinned requirements, experiment protocols and corrections, count-level and replicate-level sufficient statistics, policy decisions, CSV summaries, figures, configuration and seeds, SHA-256 receipts, and independent numerical-verification scripts and reports. Corrected deterministic run-v3 and identifier-study/run-v1 are result-bearing. Historical failed runs and their correction records are clearly labeled and retained for transparency. The data are synthetic; reproduction requires no external dataset or model API. Research scripts and their usage documentation are licensed under MIT. Generated numerical records, saved log statistics, decision vectors, result tables, figures and experiment protocols are licensed under Creative Commons Attribution 4.0 International. These are material-specific licenses, not a choice of either license for every file. LICENSING.md and the supplied license texts define the scopes. Third-party materials retain their original licenses. This record releases the complete code and generated numerical records publicly before journal publication; the sealed archive retains its original preparation-time availability statements. Version 4 updates editorial materials, acknowledgments, and research-assistance disclosures. It preserves the revision-3 research code, fixed protocols, generated data, numerical results, and prior verification evidence; no experiment results or numerical-review reruns are added. Version 4 is the local package revision and does not imply earlier Zenodo records. This work was self-funded. Contact: aojha47@gatech.edu. Research assistance disclosure: Generative AI tools, including OpenAI Codex, were used for literature search, manuscript drafting, code implementation, numerical checking, and the development of candidate formulations and draft proof arguments. The author reviewed and revised the work and is responsible for all claims, proofs, results, and citations.