Skip to content
#small language model Open access

From Shutdown Resistance to Self-Continuation Control: Identifiability, Intervention Stability, and Evidence Standards for Artificial Agents

Oct 2026 · Zenodo (CERN European Organization for Nuclear Research) · 3 references
Ethics and Social Impacts of AI

Abstract

Recent evaluations show that frontier language-model agents can resist shutdown, override human control, protect peers, and exhibit other preservation-oriented behaviors. These findings are important for safety, but a behavioral act of resistance does not by itself identify what is being preserved or why. We develop an identification framework for distinguishing task-instrumental continuation, current-bearer self-continuation, successor or peer continuation, and other persistence targets. The framework combines exact counterexamples, path-blocking contrasts, joint-consequence analysis, and small learned-policy experiments. We first show that bundled shutdown comparisons can make direct current-bearer value, task value, and their interaction behaviorally identical. We then show that separately accurate predictions of current-bearer and successor availability can remain decision-insufficient when task value depends on their joint law. In learned synthetic systems, an evaluator-declared bearer-role predictive relation can be acquired from an initially zero pathway when prediction requires it, yet flexible policies can fit ancestry-distinguishing training data while extrapolating spurious continuation effects to held-out interventions. Minimal intervention supervision sharply reduces this underspecification but does not eliminate it outside the supervised range. Finally, we show that finite intervention evidence becomes a mechanism claim only relative to a declared intervention domain and hypothesis or regularity class. These results motivate a layered evidence standard for claims about artificial-agent self-continuation. A retrospective audit of the public ROGUE computer-use benchmark shows the practical value of the distinction: ROGUE provides strong evidence of corrigibility failures, but its shutdown target is the task VM rather than an independently identified current model bearer, and shutdown remains causally entangled with task completion. The framework therefore separates strong behavioral evidence from stronger mechanistic conclusions and does not treat continuation control as a measure of consciousness, negative valence, or fear. Archival methods preprint. Exact/model-relative results, synthetic learning experiments, and retrospective benchmark interpretation are explicitly separated. Not peer reviewed. The work does not establish AI consciousness, negative valence, or fear of death.

View source

Similar papers

#computer vision Open access Jun 2016

Software Development in Startup Companies: The Greenfield Startup Model

The results are packaged in the Greenfield Startup Model (GSM), which explains the priority of startups to release the product as quickly as possible, and the need to shorten time-to-market, by speeding up the development through low-precision engineering activities.

Carmine Giardino, Nicolò Paternoster, M. Unterkalmsteiner et al. · 178 citations · ⚡14
#computer vision Open access Oct 2016

Software Startups - A Research Agenda

Software startup companies develop innovative, software-intensive products within limited timeframes and with few resources, searching for sustainable and scalable business models.

M. Unterkalmsteiner, P. Abrahamsson, Xiaofeng Wang et al. · 157 citations · ⚡17
#machine learning Review Open access Oct 2016

“Failures” to be celebrated: an analysis of major pivots of software startups

This study conducts a case survey study based on the secondary data of the major pivots happened in 49 software startups, and demonstrates that customer need pivot is the most common among all pivot types.

Sohaib Shahid Bajwa, Xiaofeng Wang, Anh Nguyen-Duc et al. · 127 citations · ⚡15
#computer vision Review Open access May 2015

A survey study on major technical barriers affecting the decision to adopt cloud services

The comparison of adopter and non-adopter sample reveals three potential adoption inhibitor, security, data privacy, and portability, which underlines the importance of the technical and security perspectives for research investigating the adoption of technology.

Nattakarn Phaphoom, Xiaofeng Wang, S. Samuel et al. · 111 citations · ⚡8
#computer vision Open access Feb 2018

Lean Internal Startups for Software Product Innovation in Large Companies: Enablers and Inhibitors

This study investigates how Lean internal startup facilitates software product innovation in large companies and identifies its enablers and inhibitors, and shows the potential of the method-in-action framework to investigate the Lean startup approach in non-startup context.

Henry Edison, Nina M. Smørsgård, Xiaofeng Wang et al. · 78 citations · ⚡6
#computer vision Conference Sep 2010

Exploring the Sources of Waste in Kanban Software Development Projects

The application of agile software methods and more recently the integration of Lean practices contribute to the trend of continuous improvement in the software industry. One such area warranting proper empirical evidence is a project’s operational efficiency when using the Kanban method. This short paper takes a new an...

Marko Ikonen, Petri Kettunen, Nilay V. Oza et al. · 67 citations · ⚡9

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.