Artifact of Lessons to Define Testing Oracles for Hybrid Quantum–Classical Software
Abstract
HQC Runtime Test Oracles: Artifact Abstract Hybrid quantum–classical (HQC) software poses a fundamental test-oracle problem: executions are stochastic, correct outputs are often unknown in advance, and final results conceal the optimisation process that produced them. A decade of research has characterised major HQC failure modes—trainability collapse, sampling-complexity barriers, noise-induced degradation, optimisation pathologies, and theoretical hardness limits—but these insights exist as analytical results rather than executable checks. This artifact accompanies a framework that translates these failure characterisations into runtime test oracles. Each is mapped to a proxy measurable on one recorded optimisation trace and to a verdict rule with a documented threshold: FAIL for contract violations, WARN for statistical proxies, and UNDECIDED for intractable instances. Eight oracles are organised in four layers (trainability, robustness, execution consistency, and domain consistency), plus a hardness diagnostic, and attach to existing code through lightweight adapters. On an oracle-independent pool of reference runs of the Variational Quantum Eigensolver (VQE), targeted fault injection confirms that every oracle implements its predicate (900/900, with no false flag on clean runs). Blind injection, whose magnitudes ignore all thresholds, shows that detection follows the declared thresholds (1253/1800 detected). Across 320 admitted configurations of 15 open-source VQE repositories, the oracles report 156 failures and 71 trajectory anomalies, which led to 13 upstream reports. Every report that received a maintainer response was acknowledged as valid; four have been fixed, two further fixes have been promised, and none has been disputed. The package also contains a leave-one-out ablation, a sensitivity analysis of every threshold, and a QAOA case study showing that the domain-agnostic oracles transfer unchanged. Contents hqc_oracle_framework/ hqc_oracles/: oracle framework, eight oracles, and 59 unit tests in tests/ campaign_run/adapters/: adapters for 15 repositories examples/: RQ1 benchmarks, sensitivity analysis, and QAOA study campaign_run/results/*.jsonl: recorded campaign corpus campaign_run/revision_analysis.py: re-analysis of the campaign corpus issue_reports/repros/: upstream issue reproducers campaign_run/verify_paper_claims.py: verifier for the submitted numbers verify_revision_claims.py: verifier for the revised numbers README.md: layout, architecture, and per-repository tables REVISION.md: mapping from reviewer points to changes, files, and results CHANGELOG.md: provenance of experiments and corpus Setup Python ≥ 3.10 is required. Install the dependencies listed in requirements.txt: qiskit>=2.0 qiskit-algorithms==0.4.0 openfermion openfermionpyscf pyscf scipy numpy matplotlib pytest for the test suite Per-repository environment notes are provided in campaign_run/INSTALL.md. RQ1 — Implementation Validity and Detection Power Run: cd hqc_oracle_framework python3 examples/10_rq1_fault_injection.py --mode targeted # RQ1a python3 examples/10_rq1_fault_injection.py --mode blind # RQ1b python3 examples/11_sensitivity_analysis.py The script sweeps the H₂/STO-3G dissociation curve across six bond lengths, two circuit depths, and three seeds. A run is a control if and only if its energy is within 1.6 mHa of FCI; no oracle verdict is consulted when defining the control set. This gives 31–33 of the 36 runs, depending on the linear-algebra backend. RQ1a — Targeted injection RQ1a injects 100 instances of each of nine fault classes: 900/900 faults are detected. The per-class Wilson 95% lower bound is 96.3%. 0/31 controls are flagged. RQ1b — Blind injection RQ1b applies the same mechanisms with 200 instances per class, using magnitudes sampled log-uniformly over oracle-agnostic ranges and uniformly selected injection sites: 1253/1800 faults are detected. A single magnitude cut explains 87.5–100% of verdicts per class. 0/31 controls are flagged. Outputs are written to: examples/rq1_*_results.{json,csv} examples/visuals/ examples/latex/ The complete revision takes approximately 20 minutes on one core. Under a time limit, run individual classes using --classes ... and then combine them with --merge. The submitted benchmark, 09_statistical_fault_injection.py, is retained for traceability. RQ2 — Faults in Real-World HQC Software Run: cd hqc_oracle_framework/campaign_run python3 verify_paper_claims.py # submitted verdicts python3 revision_analysis.py # revised verdicts and analyses Submitted verdicts The original campaign contains 331 adapter/configuration pairs: 11 are ADAPTER_FAILED. 320 are admitted: 134 PASS 157 FAIL 29 warning-only Performance appears in 156 of the 157 FAILs. The G2 reproducibility audit covers 21 paired-seed runs: 17 are bit-identical. 3 differ across invocations. Revised verdicts The revised results are re-derived from the same recorded traces without re-execution. One PhysicsConsistency FAIL caused by a framework caching defect is removed; the defect is now fixed and covered by a regression test. Trajectory monotonicity is also reported separately from hard invariants. The revised corpus contains: 156 FAIL 1 warning-only 163 PASS 71 trajectory anomalies, including 20 in gradient-based exact runs The analysis script also produces: threshold-sensitivity analysis; non-chemistry tolerance analysis; post-hoc StatisticalQuery analysis; metamorphic-relation analysis. Re-running the campaign from source is optional and takes several hours. To do so: cd hqc_oracle_framework/campaign_run ./clone_repos.sh Then run the campaign driver and re-run the verifiers. No third-party source code is modified. RQ3 — What Each Oracle Contributes revision_analysis.py computes a leave-one-out ablation over the 320 admitted records. Performance dominates detection: without Performance, 150 records lose their only FAIL. Of the 185 flagged records: 100 are flagged by Performance alone; 28 by PhysicsConsistency alone (trajectory anomalies); 1 by NoiseRobustness alone; 56 by several oracles. The smallest subset reproducing all findings is {Performance, PhysicsConsistency, NoiseRobustness}. Beyond detection, process-level oracles diagnose causes including expressibility ceilings, fragile optima, and insufficient shot budgets. GradientOptimization ran on only 4 records. ExecutionConsistency and TheoreticalHardnessLimits never fire in this corpus; RQ1 nevertheless shows that each detects its target fault. Transfer Beyond VQE The unchanged oracle suite can also be run on QAOA/MaxCut: python3 examples/12_transfer_qaoa.py The experiment produces: 0/84 false flags on clean runs; 149/210 detections under blind injection. Verification Run the complete test suite with: python3 -m pytest -q tests This runs all 59 tests. To verify every revised numerical claim: python3 verify_revision_claims.py The command exits with status 0 when all revised numbers are reproduced. Upstream Issue Reports The package contains the 13 issues submitted to maintainers. Each includes a runnable minimal reproducer checked against the recorded campaign corpus: hqc_oracle_framework/issue_reports/repros/ These reproducers allow the reported findings to be independently inspected without modifying third-party source code.