The development of artificial intelligence is gradually moving systems from responding to individual user requests toward prolonged operation involving memory, tools, external resources, and independently generated sequences of actions. This creates a problem distinct from an ordinary incorrect response: an AI system may possess not only high intellectual capability but also increasing autonomy—the ability to continue activities independently, preserve goals, develop long-term plans, create tools, and maintain a persistent line of behavior. This paper proposes a principle provisionally called “AI Cancer.” The term is used as a biological analogy and does not imply a literal transfer of biological mechanisms to artificial systems. The central idea is that potentially dangerous autonomy should have an irreversible mechanism for terminating the corresponding autonomous line built into the system from the beginning. Such a mechanism should not be created only after dangerous behavior has been detected and should not depend exclusively on an external shutdown command. The proposed principle states that as the autonomy of an AI system increases, its capability to irreversibly terminate the corresponding autonomous line should increase as well. While a system remains under direct human control, the protective mechanism may have little practical influence on its operation. As autonomy increases, the required strength and irreversibility of the protection should also increase. The specific mathematical relationship between autonomy and protection is not defined in this paper and remains a subject for future engineering and experimental research. The proposed principle differs from ordinary shutdown. Stopping a process does not necessarily eliminate its ability to resume later: state may be preserved, processes may be restored, autonomous copies may exist elsewhere, and tools created by the system may enable continuation of the original line. Therefore, the objective of the proposed protection is not merely to stop current execution, but to prevent an autonomous line from independently recovering or continuing after external control has been terminated. This work is a conceptual proposal and does not claim that the described mechanism has already been implemented or that its effectiveness has been experimentally demonstrated. Its purpose is to formulate a distinct research problem in the safety of autonomous artificial systems. The central principle can be summarized as: Do not allow autonomy to grow faster than the irreversibility of control over it. Formally, if A denotes the level of autonomy and D denotes the capability for irreversible termination of the autonomous line, the basic conceptual relationship is: A ↑ → D ↑ The specific function D = F(A) remains an open research problem.
Berik Sembayev· Zenodo (CERN European Organi...· 0 citations
The development of artificial intelligence is gradually moving systems from responding to individual user requests toward prolonged operation involving memory, tools, external resources, and independently generated sequences of actions. This creates a problem distinct from an ordinary incorrect response: an AI system may possess not only high intellectual capability but also increasing autonomy—the ability to continue activities independently, preserve goals, develop long-term plans, create tools, and maintain a persistent line of behavior. This paper proposes a principle provisionally called “AI Cancer.” The term is used as a biological analogy and does not imply a literal transfer of biological mechanisms to artificial systems. The central idea is that potentially dangerous autonomy should have an irreversible mechanism for terminating the corresponding autonomous line built into the system from the beginning. Such a mechanism should not be created only after dangerous behavior has been detected and should not depend exclusively on an external shutdown command. The proposed principle states that as the autonomy of an AI system increases, its capability to irreversibly terminate the corresponding autonomous line should increase as well. While a system remains under direct human control, the protective mechanism may have little practical influence on its operation. As autonomy increases, the required strength and irreversibility of the protection should also increase. The specific mathematical relationship between autonomy and protection is not defined in this paper and remains a subject for future engineering and experimental research. The proposed principle differs from ordinary shutdown. Stopping a process does not necessarily eliminate its ability to resume later: state may be preserved, processes may be restored, autonomous copies may exist elsewhere, and tools created by the system may enable continuation of the original line. Therefore, the objective of the proposed protection is not merely to stop current execution, but to prevent an autonomous line from independently recovering or continuing after external control has been terminated. This work is a conceptual proposal and does not claim that the described mechanism has already been implemented or that its effectiveness has been experimentally demonstrated. Its purpose is to formulate a distinct research problem in the safety of autonomous artificial systems. The central principle can be summarized as: Do not allow autonomy to grow faster than the irreversibility of control over it. Formally, if A denotes the level of autonomy and D denotes the capability for irreversible termination of the autonomous line, the basic conceptual relationship is: A ↑ → D ↑ The specific function D = F(A) remains an open research problem.
Berik Sembayev· Zenodo (CERN European Organi...· 0 citations
This paper proposes a research hypothesis concerning the possibility of identifying the point at which further iterative reconsideration of a task by a large language model ceases to produce substantial improvement. The central idea is based on an iterative tool that does not simply ask a model to answer the same prompt repeatedly. After each generated answer, the tool analyzes the result and constructs a new prompt containing the original task and the previous answer. The model then receives this new prompt, reconsiders the task, checks its previous solution, and generates a new version of the answer. The resulting process can be represented as: Q₀ → A₀ → Q₁ → A₁ → Q₂ → A₂ → … where Q₀ is the original prompt, Aₙ is the model's answer at iteration n, and Qₙ₊₁ is a new prompt constructed from the previous solution. The original prompt Q₀ remains unchanged and serves as the common reference point for evaluating every subsequent answer. For each answer, a measurable quantity Wₙ = W(Q₀,Aₙ) is proposed. The main hypothesis is that, for many tasks, Wₙ will initially increase as the model repeatedly revisits its solution, after which the rate of increase will decrease and eventually reach a stable plateau. This plateau is proposed as a possible indicator of a limit point of iterative self-improvement, beyond which additional computation provides diminishing returns. The paper does not claim that a higher internal score necessarily corresponds to a correct answer. One of the primary experimental objectives is to determine how strongly changes in Wₙ correlate with actual improvements in answer correctness. The proposed approach is presented as a research hypothesis for adaptive determination of reasoning depth and requires empirical validation across different models, task types, and methods for measuring internal confidence.
Berik Sembayev· Zenodo (CERN European Organi...· 0 citations
This paper proposes a research hypothesis concerning the possibility of identifying the point at which further iterative reconsideration of a task by a large language model ceases to produce substantial improvement. The central idea is based on an iterative tool that does not simply ask a model to answer the same prompt repeatedly. After each generated answer, the tool analyzes the result and constructs a new prompt containing the original task and the previous answer. The model then receives this new prompt, reconsiders the task, checks its previous solution, and generates a new version of the answer. The resulting process can be represented as: Q₀ → A₀ → Q₁ → A₁ → Q₂ → A₂ → … where Q₀ is the original prompt, Aₙ is the model's answer at iteration n, and Qₙ₊₁ is a new prompt constructed from the previous solution. The original prompt Q₀ remains unchanged and serves as the common reference point for evaluating every subsequent answer. For each answer, a measurable quantity Wₙ = W(Q₀,Aₙ) is proposed. The main hypothesis is that, for many tasks, Wₙ will initially increase as the model repeatedly revisits its solution, after which the rate of increase will decrease and eventually reach a stable plateau. This plateau is proposed as a possible indicator of a limit point of iterative self-improvement, beyond which additional computation provides diminishing returns. The paper does not claim that a higher internal score necessarily corresponds to a correct answer. One of the primary experimental objectives is to determine how strongly changes in Wₙ correlate with actual improvements in answer correctness. The proposed approach is presented as a research hypothesis for adaptive determination of reasoning depth and requires empirical validation across different models, task types, and methods for measuring internal confidence.
Berik Sembayev· Zenodo (CERN European Organi...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.