Skip to content
#edge computing Open access

Preregistration: Forecasting Crises Inside Autocracies with Historical Backcasts

Sep 2026 · Zenodo (CERN European Organization for Nuclear Research)

Abstract

Preregistered protocol, version 1.4, for sprint 4 of seldon-lab, the research workflow of the Computational Macrohistory (CMH) programme at the Foundations Institute of Computational Social Science (FICSS). Question. Among non-democracies (V-Dem polyarchy below 0.5), does a model with slow covariates forecast crises better than the persistence of each state's own events and better than the base rate of autocracies, and at which horizons (2, 5, 10 and 30 years)? Is the democracy level related to the hazard as an inverted U? Events. The primary confirmatory event is the irregular exit of the effective leader (Archigos) with no irregular regime end of the same state in V-Dem within two years; the secondary events are successful coups (Powell and Thyne) and autocratic breakdowns (Geddes, Wright and Frantz), under the same rule. An onset within five years of the last kept onset of the same state belongs to the same crisis. Design. Logistic models fitted directly at each horizon and refitted at each origin on the past only (rolling-origin backcasts). The confirmatory model adds the democracy level (as an inverted U) and log GDP per capita to a block of the state's own event history; the reference of the closure test is a persistence model on that history alone. Forecasts are scored with the Brier skill score against the persistence model and against the base rate of autocracies, with conservative bootstrap lower bounds over states and over decades. The closure test is a 2x2 table: a success requires the model to beat both the persistence model and the base rate at 10 years or more, and the ledger label of each of the four outcomes is fixed in advance. The hypotheses, functional forms and expected signs were fixed before the event data were downloaded. The choice of the core covariates rests in part on results already seen on another event, the V-Dem irregular regime ends of the lab's previous sprint; for this reason only events absent from that series are confirmatory. Power. The minimum effect of interest is a Brier skill score of 0.05 at 10 years against the persistence model, set by convention. With the minimum of 180 confirmatory events, the power estimated in vitro is 0.75-0.80 against the persistence model and 0.59-0.78 for the success rule; the protocol declares the minimum, 0.59, as the power of the design, and reads a negative result through the minimum detectable effect. A fallback ladder that depends only on event counts is fixed in advance, so the closure test cannot end as "not evaluable", and any conclusion names the event definition it holds for. Freeze and versions. Version 1.0 was frozen on 27 September 2026 by commit a6ab5479c33148a4bcaf56af8365305ded567d4e of the lab repository, which is private, and deposited as the first version of this record (DOI 10.5281/zenodo.22992168). The new event data (Archigos, Powell and Thyne, GWF) were downloaded only after that deposit. Version 1.1 (DOI 10.5281/zenodo.23009979, commit 310c2328f6338acfc1b554bd189d8bc3ad5d902b) added Appendix A, fixed from the codebooks before any count, and three entries of Appendix B on rules that the frozen text leaves open. Version 1.2 (DOI 10.5281/zenodo.23033295, commit 34194e78093088c1559e297e0026f8737ed956f3) added the rules of Appendix A9 and the counts of Appendix A8, made before any model was fitted on the new event data: no step of the fallback ladder reaches the minimum of 180 events, so HA1 is run on the design of step 3, the irregular exits of the leader or successful coups from 1950, absent from E1, with verification origins from 1946 (the first is 1950), on N = 46 events, of which at most 43 can fall in a target at 10 years; at that N the effect detected with power 0.8 against the persistence model, extrapolated from the in-vitro estimates, is a Brier skill score of 0.101 at 10 years. Version 1.3 (commit 2c4c546, not deposited on its own) added Appendix B4: the readings of the verification code on points that the frozen text leaves open, four departures from the letter, marked (the GDP series of the sprint 3 model in HA3, the pairs of an exploratory model, the history windows in calendar years, and the coverage of the annual pairs of HA2 and HA4), five errors of the text, and the procedure of the verification run. Version 1.3 was committed before any development run on the new data. On the same day, after the development runs and before the rehearsal with synthetic outcomes, its item 23, on an aborted run, was amended and its item 24, on that rehearsal, was added; version 1.4 amends items 22 to 24 and adds item 25 after a review gate of the verification code: the marker of a used run is written after the counts of Appendix A8 are checked, the files of this deposit are committed before the verification tag, the one real series after 1945 that the rehearsal computes is named, and two edge cases of the code are declared. None changes a count, the number of events, the step of the ladder, a threshold, a decision rule or a horizon of HA1, nor the features of any model; one departure fixes how the history features of the persistence model and of the confirmatory model are counted for states with gaps in their record, and one changes the GDP series of the sprint 3 model, the comparator of HA3. Version 1.4 was committed on 2026-09-30 by commit b991ce85bda44e55493075dafacdc53b7df210e3. The body is unchanged, as the file changes-v1.0-to-v1.4.diff shows; changes-v1.2-to-v1.4.diff shows the changes from version 1.2. SHA-256 of the protocol file: 3be423914c55c5b3fb38b8f9e243ef6ec5e1fa750efa76fa5146cc81beb7fb40. This is the last deposit before the verification run: it holds the verification code, frozen after a review gate, and the run comes after it, at the verification tag of the lab repository, with exactly the code files deposited here. Files. The protocol, in Markdown and in a PDF copy for reading; the changes from versions 1.0, 1.1 and 1.2, as diffs; the in-vitro simulations of power, of the skill ceiling, of the persistence reference, of the full and core models and of the design of version 1.0 (Python, synthetic data only), with the modules they import; their output tables (CSV); the scripts that read the codes of the new event files (Appendix A7), made the counts of Appendix A8 and checked the country codes (Appendix B), with their outputs, which are totals and lists of codes and years only (these scripts need the event files, V-Dem and the official lists of states, which are not included); the specification of the verification code, which Appendix B4 cites, and its completeness critic; the verification code itself, frozen before the verification run, with its tests on synthetic data and the summaries of the development runs; and a README with the requirements (Python 3.12 or later, numpy, pandas, scipy, numba) and the commands to run the in-vitro scripts from the deposit folder. The PDF file is generated from the Markdown file without changes to the text; the Markdown file, with the SHA-256 given above, is the reference. No third-party data are included. Prospective forecasts for 2026-2035 made under this protocol will be deposited separately. Use of AI tools. I used AI tools in this work. Claude models (Anthropic) helped me write and check the code, run the simulations, and draft and revise the text. DeepSeek models did smaller tasks, such as literature searches and first checks. All tasks and checks are logged in the laboratory records. I checked the results and made all research choices. I am responsible for the content.

View source

Similar papers

#computer vision Review Sep 2017

Agile Software Development Methods: Review and Analysis

This publication proposes a definition and a classification of agile software development approaches and analyses ten software development methods that can be characterized as being "agile" against the defined criterion.

P. Abrahamsson, O. Salo, Jussi Ronkainen et al. · 727 citations · ⚡54
#computer vision Jun 2008

The impact of agile practices on communication in software development

The study shows that agile practices improve both informal and formal communication, but indicates that, in larger development situations involving multiple external stakeholders, a mismatch of adequate communication mechanisms can sometimes even hinder the communication.

M. Pikkarainen, Jukka Haikara, O. Salo et al. · 401 citations · ⚡48
#machine learning Review Open access Oct 2014

Software development in startup companies: A systematic mapping study

The results indicate that software engineering work practices are chosen opportunistically, adapted and configured to provide value under the constrains imposed by the startup context.

Nicolò Paternoster, Carmine Giardino, M. Unterkalmsteiner et al. · 394 citations · ⚡54

Related blog posts

Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.