Skip to content
#software testing Open access

Artifact of Lessons to Define Testing Oracles for Hybrid Quantum–Classical Software

Sep 2026 · Zenodo (CERN European Organization for Nuclear Research)

Abstract

HQC Runtime Test Oracles: Artifact Abstract Hybrid quantum–classical (HQC) software poses a fundamental test-oracle problem: executions are stochastic, correct outputs are often unknown in advance, and final results conceal the optimisation process that produced them. A decade of research has characterised major HQC failure modes—trainability collapse, sampling-complexity barriers, noise-induced degradation, optimisation pathologies, and theoretical hardness limits—but these insights exist as analytical results rather than executable checks. This artifact accompanies a framework that translates these failure characterisations into runtime test oracles. Each is mapped to a proxy measurable on one recorded optimisation trace and to a verdict rule with a documented threshold: FAIL for contract violations, WARN for statistical proxies, and UNDECIDED for intractable instances. Eight oracles are organised in four layers (trainability, robustness, execution consistency, and domain consistency), plus a hardness diagnostic, and attach to existing code through lightweight adapters. On an oracle-independent pool of reference runs of the Variational Quantum Eigensolver (VQE), targeted fault injection confirms that every oracle implements its predicate (900/900, with no false flag on clean runs). Blind injection, whose magnitudes ignore all thresholds, shows that detection follows the declared thresholds (1253/1800 detected). Across 320 admitted configurations of 15 open-source VQE repositories, the oracles report 156 failures and 71 trajectory anomalies, which led to 13 upstream reports. Every report that received a maintainer response was acknowledged as valid; four have been fixed, two further fixes have been promised, and none has been disputed. The package also contains a leave-one-out ablation, a sensitivity analysis of every threshold, and a QAOA case study showing that the domain-agnostic oracles transfer unchanged. Contents hqc_oracle_framework/ hqc_oracles/: oracle framework, eight oracles, and 59 unit tests in tests/ campaign_run/adapters/: adapters for 15 repositories examples/: RQ1 benchmarks, sensitivity analysis, and QAOA study campaign_run/results/*.jsonl: recorded campaign corpus campaign_run/revision_analysis.py: re-analysis of the campaign corpus issue_reports/repros/: upstream issue reproducers campaign_run/verify_paper_claims.py: verifier for the submitted numbers verify_revision_claims.py: verifier for the revised numbers README.md: layout, architecture, and per-repository tables REVISION.md: mapping from reviewer points to changes, files, and results CHANGELOG.md: provenance of experiments and corpus Setup Python ≥ 3.10 is required. Install the dependencies listed in requirements.txt: qiskit>=2.0 qiskit-algorithms==0.4.0 openfermion openfermionpyscf pyscf scipy numpy matplotlib pytest for the test suite Per-repository environment notes are provided in campaign_run/INSTALL.md. RQ1 — Implementation Validity and Detection Power Run: cd hqc_oracle_framework python3 examples/10_rq1_fault_injection.py --mode targeted # RQ1a python3 examples/10_rq1_fault_injection.py --mode blind # RQ1b python3 examples/11_sensitivity_analysis.py The script sweeps the H₂/STO-3G dissociation curve across six bond lengths, two circuit depths, and three seeds. A run is a control if and only if its energy is within 1.6 mHa of FCI; no oracle verdict is consulted when defining the control set. This gives 31–33 of the 36 runs, depending on the linear-algebra backend. RQ1a — Targeted injection RQ1a injects 100 instances of each of nine fault classes: 900/900 faults are detected. The per-class Wilson 95% lower bound is 96.3%. 0/31 controls are flagged. RQ1b — Blind injection RQ1b applies the same mechanisms with 200 instances per class, using magnitudes sampled log-uniformly over oracle-agnostic ranges and uniformly selected injection sites: 1253/1800 faults are detected. A single magnitude cut explains 87.5–100% of verdicts per class. 0/31 controls are flagged. Outputs are written to: examples/rq1_*_results.{json,csv} examples/visuals/ examples/latex/ The complete revision takes approximately 20 minutes on one core. Under a time limit, run individual classes using --classes ... and then combine them with --merge. The submitted benchmark, 09_statistical_fault_injection.py, is retained for traceability. RQ2 — Faults in Real-World HQC Software Run: cd hqc_oracle_framework/campaign_run python3 verify_paper_claims.py # submitted verdicts python3 revision_analysis.py # revised verdicts and analyses Submitted verdicts The original campaign contains 331 adapter/configuration pairs: 11 are ADAPTER_FAILED. 320 are admitted: 134 PASS 157 FAIL 29 warning-only Performance appears in 156 of the 157 FAILs. The G2 reproducibility audit covers 21 paired-seed runs: 17 are bit-identical. 3 differ across invocations. Revised verdicts The revised results are re-derived from the same recorded traces without re-execution. One PhysicsConsistency FAIL caused by a framework caching defect is removed; the defect is now fixed and covered by a regression test. Trajectory monotonicity is also reported separately from hard invariants. The revised corpus contains: 156 FAIL 1 warning-only 163 PASS 71 trajectory anomalies, including 20 in gradient-based exact runs The analysis script also produces: threshold-sensitivity analysis; non-chemistry tolerance analysis; post-hoc StatisticalQuery analysis; metamorphic-relation analysis. Re-running the campaign from source is optional and takes several hours. To do so: cd hqc_oracle_framework/campaign_run ./clone_repos.sh Then run the campaign driver and re-run the verifiers. No third-party source code is modified. RQ3 — What Each Oracle Contributes revision_analysis.py computes a leave-one-out ablation over the 320 admitted records. Performance dominates detection: without Performance, 150 records lose their only FAIL. Of the 185 flagged records: 100 are flagged by Performance alone; 28 by PhysicsConsistency alone (trajectory anomalies); 1 by NoiseRobustness alone; 56 by several oracles. The smallest subset reproducing all findings is {Performance, PhysicsConsistency, NoiseRobustness}. Beyond detection, process-level oracles diagnose causes including expressibility ceilings, fragile optima, and insufficient shot budgets. GradientOptimization ran on only 4 records. ExecutionConsistency and TheoreticalHardnessLimits never fire in this corpus; RQ1 nevertheless shows that each detects its target fault. Transfer Beyond VQE The unchanged oracle suite can also be run on QAOA/MaxCut: python3 examples/12_transfer_qaoa.py The experiment produces: 0/84 false flags on clean runs; 149/210 detections under blind injection. Verification Run the complete test suite with: python3 -m pytest -q tests This runs all 59 tests. To verify every revised numerical claim: python3 verify_revision_claims.py The command exits with status 0 when all revised numbers are reproduced. Upstream Issue Reports The package contains the 13 issues submitted to maintainers. Each includes a runnable minimal reproducer checked against the recorded campaign corpus: hqc_oracle_framework/issue_reports/repros/ These reproducers allow the reported findings to be independently inspected without modifying third-party source code.

View source

Similar papers

#computer vision Review Sep 2017

Agile Software Development Methods: Review and Analysis

This publication proposes a definition and a classification of agile software development approaches and analyses ten software development methods that can be characterized as being "agile" against the defined criterion.

P. Abrahamsson, O. Salo, Jussi Ronkainen et al. · 727 citations · ⚡54
#computer vision Jun 2008

The impact of agile practices on communication in software development

The study shows that agile practices improve both informal and formal communication, but indicates that, in larger development situations involving multiple external stakeholders, a mismatch of adequate communication mechanisms can sometimes even hinder the communication.

M. Pikkarainen, Jukka Haikara, O. Salo et al. · 401 citations · ⚡48
#machine learning Review Open access Oct 2014

Software development in startup companies: A systematic mapping study

The results indicate that software engineering work practices are chosen opportunistically, adapted and configured to provide value under the constrains imposed by the startup context.

Nicolò Paternoster, Carmine Giardino, M. Unterkalmsteiner et al. · 394 citations · ⚡54
#computer vision Review Mar 2008

Agile methods in European embedded software development organisations: a survey on the actual use and usefulness of Extreme Programming and Scrum

The results show that the embedded industry has been able to apply agile methods in its development processes and that the appreciation of the agile methods and their individual practices appears to increase once adopted and applied in practice.

O. Salo, P. Abrahamsson · 238 citations · ⚡9
#computer vision Open access Jul 2017

What happens when software developers are (un)happy

Consequences of happiness and unhappiness that are beneficial and detrimental for developers' mental well-being, the software development process, and the produced artifacts are found.

D. Graziotin, Fabian Fagerholm, Xiaofeng Wang et al. · 236 citations · ⚡13
#computer vision Open access Oct 2004

Mobile-D: an agile approach for mobile application development

The Mobile-D approach is briefly outlined here and the experiences gained from four case studies are discussed, which helped develop an agile development approach for mobile application development.

P. Abrahamsson, Antti Hanhineva, H. Hulkko et al. · 225 citations · ⚡18

Related blog posts

MIT News · Artificial Intelligence Oct 2, 2026

Documenting the tech worker movement

Writing as a participant and researcher, PhD student JS Tan SM ’22 has co-authored a new book about the rise of tech worker protests and the employer backlash that followed.

GPT-Lab Sep 23, 2026

Requirements Don’t Live in Isolation: What We’re Exploring with Req-Space

Requirements in large systems rarely exist in isolation. Their meaning depends on the wider project context - other requirements, policies, decisions, tests, and implementation details. That becomes especially important when AI is used for review, because spotting a possible conflict or gap is only the beginning. ReqSpace explores how AI, visualisation, and connected project context can help reviewers understand those findings, trace the relationships behind them, and focus on the questions that…

GPT-Lab Sep 17, 2026

Beyond Prompt Engineering: The Role of Tacit Knowledge in Software Engineering

AI is making software generation faster, but speed does not remove the need for expertise. As more work is delegated to AI, tacit knowledge may become one of the most important human advantages in software engineering. The post Beyond Prompt Engineering: The Role of Tacit Knowledge in Software Engineering appeared first on GPT-Lab.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.