Warrant study: tasks, harness and run records for an exploratory evaluation of evidence-gated acceptance of code written by an AI agent
Materials for an exploratory evaluation of Warrant (https://doi.org/10.5281/zenodo.23029803). Twenty small Python command-line tasks, each with plain-language promises, visible checks, hidden checks, an audit suite, a reference implementation and deliberately broken implementations; the validity gates; the harness; Par...