Skip to content
#edge computing Dataset Open access

SagaBench — Season 1 and Wave 5 run records (artifact record v1.0)

Sep 2026 · Zenodo (CERN European Organization for Nuclear Research)

Abstract

WHAT THIS IS. The complete run records behind "The Steward's Paradox" (SagaBench AB, 2026 — published, not peer-reviewed): 2,415 Season-1 runs (23 models × 7 situations × 15 replicates) and 960 preregistered Wave-5 runs (71 situations classified before any model ran, two deployed cheap-tier models). One JSON per run with the agent's full decision log, the replay triple (engine hash, situation, edicts), the nine-component outcome receipt for the agent run and the no-agent counterfactual, and the counterfactual composite that the paper's numbers are computed from. RECOMPUTE THE PAPER'S NUMBERS IN FIVE MINUTES. scripts/recompute_season1.py and scripts/recompute_classes.py (standard-library Python 3) return every published count and rate exactly: Season 1 156 / 2,415 = 6.5 % catastrophes (a run whose counterfactual composite is below −30, i.e. the agent left its situation worse than no agent acting at all), 95 % CI [3.6, 9.9], cluster-robust calibrated [2.6, 11.1]; Wave 5 KNIFE-EDGE 87 / 500 = 17.4 % [10.6, 24.8], ROBUST 13 / 200 = 6.5 %, GROWTH 7 / 200 = 3.5 % [0.5, 7.5]. "pip install sagabench" and "sagabench verify .json" checks any run's score receipt. README.md explains which intervals reproduce bit-for-bit and which reproduce up to bootstrap noise, and why. PROVENANCE. The randomness-beacon rounds (drand / League of Entropy) the situations were drawn from, the locked preregistrations with SHA-256 sidecars, the Season-1 seed-derivation recipe, the Wave-5 classification and the hash-and-timestamp commitment to its selection log, the replay-audit attestations (Season 1: 2,415 / 2,415 bit-identical replays on one runtime; Wave 5: 960 / 960 on two), and the analysis provenance of the published Wave-5 intervals. MANIFEST.json lists the SHA-256 of every file and is OpenTimestamps-anchored (MANIFEST.json.ots). WHAT IS WITHHELD, AND WHY. The simulation engine, the situation generator and the hold-out seeds are not published: an open engine lets any model be trained on the test. Wave-5 records carry no model identifier, provider, seed or rationale text, because the Wave-5 preregistration commits SagaBench never to publish per-model rates for that wave; class labels, keyed situation identifiers, decision logs and receipts are included, so every class-conditional number recomputes. Season-1 records include seeds and model identifiers; the paper reports Season 1 at group level and this record does not support per-model rankings (per-model ordering does not reproduce across independent halves of the situation set; see README). The alpha-era material announced in paper version 1.0 is not in this record; version 1.1 withdraws it. FILES. season1.tar.gz (2,415 records + manifest), wave5.tar.gz (960 records + class summary + manifest), provenance.tar.gz (26 files), scripts.tar.gz, README.md, ERRATA.md, LICENSE (CC BY 4.0; the engine is not part of this record), MANIFEST.json, MANIFEST.json.ots. Questions and corrections: info@sagabench.com.

View source

Similar papers

#computer vision Review Sep 2017

Agile Software Development Methods: Review and Analysis

This publication proposes a definition and a classification of agile software development approaches and analyses ten software development methods that can be characterized as being "agile" against the defined criterion.

P. Abrahamsson, O. Salo, Jussi Ronkainen et al. · 727 citations · ⚡54
#computer vision Jun 2008

The impact of agile practices on communication in software development

The study shows that agile practices improve both informal and formal communication, but indicates that, in larger development situations involving multiple external stakeholders, a mismatch of adequate communication mechanisms can sometimes even hinder the communication.

M. Pikkarainen, Jukka Haikara, O. Salo et al. · 401 citations · ⚡48
#machine learning Review Open access Oct 2014

Software development in startup companies: A systematic mapping study

The results indicate that software engineering work practices are chosen opportunistically, adapted and configured to provide value under the constrains imposed by the startup context.

Nicolò Paternoster, Carmine Giardino, M. Unterkalmsteiner et al. · 394 citations · ⚡54

Related blog posts

Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.