Result data for: Can Large Language Model and Vision-Language Model Judges Enhance Automated Canary Analysis? A Per-Family Evaluation Against Kayenta's Statistical Judge
Abstract
Every result artefact behind the paper, in the directory layout the analysis code expects, together with the scripts that recompute the paper's tables from them and the analysis code that turns them into every number the article reports. The testbed that produced the measurements is the kayenta-ai-canary-judge repository, archived separately at 10.5281/zenodo.22553878, where all of its versions are listed. Most of the scripts here import their metric implementations from it, so a regenerated table is produced by the same code as the published one. Which directories to use. The paper's numbers come from results/corrected/ (the 180-scenario benchmark), results/pilot-45-corrected/ (the 45-scenario pilot) and results/n5-corrected/ (the five-seed sweep). The runs under reference-results/ are the earlier, uncorrected ones; they are retained as evidence, not for use. Pre-correction files that sat in results/ until version 1.0.0 are in superseded-pre-correction/. README.md opens with a table stating which directory is which. Why the uncorrected runs are here. The paper reports a defect in the evaluation harness: a tolerant parser turned a failed model call into a well-formed failing verdict at score zero with an empty error field, indistinguishable from a genuine judgement. The affected configurations were re-run in full under the fixed harness, and every published number comes from that re-measurement. The earlier runs are kept beside the corrected ones so that the reported defect can be checked independently rather than taken on trust. PROVENANCE.md documents what the defect was, how it was found, and exactly which rows it touched. Version 1.1.0. The corrective work released here moves the pre-correction files out of results/, adds the post-correction tables and the multi-seed interval figure that version 1.0.0 lacked, fixes the regeneration tools and corrects the documentation. It was prepared as 1.0.1; the minor number moves because the release is no longer corrections alone. It also adds analysis/ — the code that turns these results into every number, table and figure the article reports, together with the FACTS.md file it writes, one line per number — and two new experiments, results/name-attribution-r1/ and results/name-attribution-r2-hosted/. No reported value changes. CHANGELOG.md lists every change and what a reader of 1.0.0 could have been misled on. Two things to know before running analysis/. It is pinned to two exact versions: this record, and software 1.0.0 — not the latest. The analysis audits the instrument as 1.0.0 had it, quoting thirty-three of its source lines by file and line number and re-running two of its decision functions to reproduce a defect the corrective release fixes, so against a corrected tree it stops by design; analysis/README.md section 1 gives the two one-line checks that confirm you have the right pair. It also needs the testbed stack up: one area audits this record by re-running its own regeneration scripts, two of which render inside the canaryllm-judge-service image and one of which calls the Kayenta service over HTTP, so a handful of facts are properties of the machine as well as of the archive and analysis/README.md section 6 names them. The fact files shipped here are what this archive's own scripts produce from it, with one stated exception: after the six area scripts, build_facts.py exits 0 with every area unchanged on two consecutive runs, and section 6 of analysis/README.md gives the SHA-256 to compare against, and records that the two scans which read the record's own text — area E's list of the file paths the record names and does not contain, and FB746, which quotes verbatim every line of the record matching “PPV”, “prevalence” or “positive predictive” — both skip the fact files, which are their own output, so that a rebuild of this record reproduces it. One input is fetched rather than shipped: the Spinnaker issue thread that one section replays is the writing of people who are not the authors here and cannot be relicensed, so analysis/README.md section 4 gives the two-line curl command that reproduces both files byte for byte, and the script stops with that command in its message if they are missing. No part of that thread is in this record. The Kayenta source files the same area reads are included, unmodified, with the Apache-2.0 LICENSE and a NOTICE recording their upstream commit and copyright. The name-attribution runs. results/name-attribution-r1/ holds three runs made for the article's first revision, in which the 180 benchmark scenarios were judged again by the same eight local models with the metric names — and, in the third arm, the order in which the metrics are presented — changed. In the reported runs those names carry a suffix that is monotone in family order, so a single threshold on it reproduces every label; the runs break that one variable at a time, and then through the released own-data anonymiser as a practitioner would use it. The statistical judge, which reads the series and never sees a name, returned 3,200 verdicts and scores identical to the reported run's across the three arms, which is the control. The third arm exists because an independent check found that the first had also reordered the metrics on the benchmark's only three-metric family; CORRECTIONS.md in that directory records that finding and two others, corrects the affected records in place, and keeps every superseded record beside its replacement. Sections 0–13 of FACTS_G.md were written before the confound was found: 86 of their lines carry an inline WITHDRAWN marker and section 19 is the register, so read analysis/area-G-name-attribution/README.md before citing an FG id. The directory carries all three input datasets, all three scenario mappings, the permutation seed, the anonymiser salt, every per-call judge payload including the chart images the vision models were shown, the console logs, and a MANIFEST.csv with the sha256 of all 5,133 files; results/name-attribution-r1/README.md states what was run, on which host, with which model weight digests, at which testbed commit and with which decoding settings. The three hosted judges were not run in that revision, for want of credentials in the run sessions, and results/name-attribution-r2-hosted/ is the re-run made for the article's second revision that closes it: the same three arms put to claude-opus-4-8-vlm, gpt-oss-120b-llm and qwen3-vl-235b-vlm, nine judge-arms, 1,140 scored calls and 1,158 preserved payload records, with no AI error row and no statistical error row and with three control checks passing on nine arms of nine. Read it with that directory's RUN_MODE.md: its rows were produced at release/v1.1.0, commit 2eb7c34, with the canary-config name and description held at the values the first revision's arms ran with, and that sentence travels with every number taken from it. The hosted name result is measured on the 160 single-variable windows, and the 20 windows of the benchmark's only three-metric family are reported separately as an order measurement rather than pooled with them. One provenance note about two figures. results/figures/family_heatmap.png and results/corrected/figures/family_heatmap.png were rendered before tools/render_figures.py changed its colormap from RdYlGn to RdYlBu for colour-blind safety. Re-rendering them from the deposited code gives the same pixel dimensions, the same cells and the same labels in different colours. The two files are the only ones in this record that the deposited code does not reproduce byte for byte; everything else it draws does. The data is CC-BY-4.0; the scripts under tools/ are Apache-2.0, as in the repository they were written for; the code under analysis/ is released with the data under CC-BY-4.0, except analysis/area-F-instrument-audit/kayenta-5310ab7/, which holds nine unmodified source files of the Spinnaker Kayenta project and stays under that project's Apache License, Version 2.0, with the full licence text in that folder's LICENSE, the attribution in its NOTICE and the upstream path of each file in its SOURCES.txt. See README.md for the layout, PROVENANCE.md for which artefacts are pre- and post-correction, tools/README.md for the table-by-table regeneration map including what cannot be regenerated, and analysis/README.md for how the article's numbers are computed.