Result data for: Can Large Language Model and Vision-Language Model Judges Enhance Automated Canary Analysis? A Per-Family Evaluation Against Kayenta's Statistical Judge
Every result artefact behind the paper, in the directory layout the analysis code expects, together with the scripts that recompute the paper's tables from them and the analysis code that turns them into every number the article reports. The testbed that produced the measurements is the kayenta-ai-canary-judge reposito...