COSPEC Comparison Harness and Evaluation Dataset
An automated comparison harness and the complete 150-trial dataset supporting the evaluation of COSPEC, Spec-Driven Development using GitHub Spec Kit, and vibe coding. The dataset contains five conditions, two cross-model Maker/Director pairs (Claude and Codex), three Maker reasoning-effort levels, and five repetitions...