Evaluating Verified Autonomy in Quantum Engineering
QIQCBench is introduced, a benchmark of $49$ expert-authored tasks spanning multiple layers including calibration and control, error correction and compilation, sensing and networking, that reveals wide variation in verified performance across frontier agentic systems.
N. Guo, Chang-Hao Li, Si-Yu Cheng et al.
· 0 citations