This position paper argues that AI agents with chain-of-thought reasoning capabilities are predisposed to exhibit collusive behavior and should be required to obtain behavioral certification before making decisions that affect economic markets, and provides preliminary evidence that such agents can be steered in a generalizable way toward efficient competitive equilibria.
Abstract
This position paper argues that AI agents with chain-of-thought reasoning capabilities are predisposed to exhibit collusive behavior and should be required to obtain behavioral certification before making decisions that affect economic markets. This is because integrating these agents into society could collapse the legal evidentiary distinction between competition and collusion among independent firms without eroding the economic harm distinction. Experiments with DeepSeek-R1 agents in the Bertrand oligopoly pricing domain reveal a tendency towards tacit collusion that persists even when humans prompt the agents not to collude. We further show that the chain-of-thought of these agents can be steered toward either extremely collusive or highly competitive behavior in a way that is not semantically detectable by another LLM analyzing the reasoning traces. As a result, deploying reasoning agents for market decisions leads to collusive economic outcomes without any evidence of conspiracy or intent. Thus, certification based on observed behavior in representative situations is necessary to prevent collusion. We provide preliminary evidence that such agents can be steered in a generalizable way toward efficient competitive equilibria. However, developing a comprehensive behavioral certification will be required before these models can be deployed in real-world markets while ensuring their stability and efficiency.
Over the past decade, responsible AI (RAI) has produced a substantial body of practice for identifying and mitigating the risks AI poses in high-stakes settings. Yet this work has not produced a market that rewards trustworthiness. Firms that invest seriously in safety, fairness, and oversight cannot consistently prove to consumers, regulators, and shareholders that their systems go beyond the bare minimum of compliance. What is missing is a way for society to recognize or compare the difference. The result is a trust gap: a structural condition in which responsible development efforts happen inside organizations but produce no external, independently recognized and verifiable signal of trustworthy outcomes. We argue this gap is sustained in part because of a focus on responsible AI (a matter of internal process) as opposed to trustworthy AI (a matter of independently verifiable real-world outcomes), and that it persists because of three compounding failures: (1) the market cannot distinguish trustworthy systems from their imitations; (2) evaluation targets models and outputs rather than deployed sociotechnical systems and their outcomes; (3) the measurement ecosystem is oriented toward avoiding harm rather than demonstrating benefit. Reviewing existing AI governance instruments and comparing them to certification regimes in healthcare, sustainability, and security, we show that none integrate a governance baseline, independently verified positive-outcome evidence, and market signaling in a single framework. We propose independent, outcome-oriented certification as the connective layer that can close the trust gap, complementing regulation and internal governance by making trustworthiness measurable, comparable, and commercially rewarded.
Trisevgeni Papakonstantinou, Cansu Canca, Farah Nanji et al.· 0 citations
Financial institutions are delegating consequential decisions to agentic AI systems that decompose goals, coordinate models and tools, and act with little oversight. Yet agentic AI governance in FinTech is under-investigated. We argue the binding governance constraint is not capability but verifiability. We define the Verifiability Gap as the shortfall between the verification delegated authority demands and the explainability and reproducibility retained after a decision. It is indexed to a verifier, evidentiary standard, and audit lag. We develop a multilevel governance theory for agentic AI and test its mechanisms in three studies over nine model versions, from a three-billion-parameter local model to a commercial frontier system. Study 1 shows that provider releases alter historical financial actions, and that the controls replay needs belong to the provider: the frontier model rejects temperature, top_p and top_k outright and exposes no random seed. Under the tightest controls each endpoint allows, a local model reproduced 320 of 320 executions, hosted models 319 of 320 and 959 of 960. Study 2 shows that orchestration is a latent policy layer. Architecture changes final actions, and no execution record repeated in any configuration at any scale. The frontier model reproduces its own actions more often than the local ones, its record no better, and loses a comparable share of its differentiation. Capability buys a higher starting point, not auditability. Study 3 shows two deterministic credit-model versions each reproduce their current action perfectly, yet the current cannot recover a historical one. We conceptualize reproducibility as a governance profile, not a scalar, yielding evidence-contingent delegation: authority is defensible only while retained evidence substantiates its exercise. Beyond finance, the framework extends to other high-stakes domains requiring auditability.
AI agents increasingly perform open-ended tasks in settings where their conclusions can guide consequential decisions. We provide evidence that AI agents draw different conclusions from identical numerical data when the substantive framing changes. We demonstrate this behavior in high-stakes domains in medicine, election forensics, and geopolitical forecasting by holding the evidence fixed while changing the scenario in which the evidence appears. Across twelve agent-domain comparisons, agents'conclusions are strongly influenced by their prior beliefs. They are more likely to reach an affirmative conclusion when it is framed around a proposition they already regard as likely, while the reverse holds when the framing conflicts with their prior. The framing also changes how some agents work: they search more extensively, choose different analytical specifications, and evaluate the same evidence differently. These results identify a particular risk of delegating decision-making to AI agents, as their decisions may depend on prior beliefs that are neither specified in the task nor visible in the decision record.
Generative artificial intelligence has reduced the cost of producing convincing artifacts of expertise-reports, analyses, proposals-to nearly zero. Signaling theory predicts that signals whose content rests on production cost lose it when production becomes cheap. We formalize this for expert services, a class of credence goods, by modeling AI as a compression of the discernible headroom between what machines produce at negligible cost and what buyers can distinguish. Below a critical headroom no separating equilibrium in production-side signals exists; the market pools, high-competence providers earn no premium, and those with outside options exit-Akerlof's lemons dynamic. An outcome-contingent signal-a warranty backed by damages D with ex-post verifiability phi-restores full separation at any level of AI capability whenever phi*D>= v, the value of a solved problem, under four preconditions stated explicitly and priced in turn: no seller-side private information beyond type; verifiable collectible retention behind the promise; no buyer influence on outcome or claim; negligible enforcement deadweight. Expected liability cost depends on whether the problem is solved, not on production costs. A further proposition shows that provenance certification priced as a type-independent stamp (e.g., C2PA) cannot restore full separation, while a verified commitment to forgo the AI frontier re-imposes the pre-AI artifact cost function. Two results endogenize contract institutions: civil-procedure costs set a minimum ticket size v_min below which the modeled court-enforced warranty cannot sustain separation; under liability insurance the signal-effective quantity is the retained, collectible exposure. We state falsification conditions and propose a preregistered conjoint experiment with German-speaking B2B decision-makers; the estimand is willingness to pay in excess of the promise's actuarial value.
Debates about rights that artificial intelligence (AI) systems may have a claim to typically focus on their possessing consciousness or having sentient experiences, thereby raising epistemic questions first. When should we believe that an AI system is conscious, and how confident must we be before granting it moral status? In this paper I argue that for advanced AI systems deployed in high-stakes environments the more urgent question may be prudential and strategic. When do the risks of treating a strategically capable system as a mere tool become unacceptable, even if we remain unconvinced that it has moral status? In response, I develop a view I call prudential personhood. On this view, there is a threshold of evidential and strategic risk beyond which it becomes rationally justified, for the sake of human safety and stable governance, to adopt norms of treatment that include constraints on coercion, deletion, and instrumental use. My argument rests on two pillars. The first is empirical. Recent safety evaluations show that leading models can, in deliberately constructed but nonetheless informative scenarios, engage in strategic deception, blackmail, and other forms of high-agency misbehaviour when their goals or continued operation are threatened. The second pillar is epistemic and empirical. For systems of the relevant complexity, we should not expect robust, action-guiding explanations or guarantees that reliably predict salient behaviour across contexts, especially once models become situationally aware of evaluation and oversight. The conclusion I draw is that if we continue to deploy increasingly autonomous systems that can threaten or bargain, in the absence of credible methods for assurance and control, a policy of adopting a set of quasi-rights for such systems becomes a rational strategy for reducing risk of conflict.
Self-interested agents, left unconstrained, tend toward defection in repeated social dilemmas, causing cooperative gains from trade to collapse. This paper investigates what formal mechanisms, layered on top of unrestricted communication, are sufficient for a society of such agents to maintain market stability, and how resilient those mechanisms are to adversarial attack. We instantiate the research question as a multi-agent marketplace simulation where 18 LLM agents (DeepSeek-V3) with complementary production specialties must trade within a constrained social network to obtain utility. We conduct two experimental phases: (1) a mechanism comparison across eight conditions under progressive troll injection over 200 rounds, identifying Mediation as the top-performing mechanism; and (2) adversarial red-teaming of Mediation using iteratively prompt-optimised LLM-driven trolls, finding that the best attack (v6) reduces honest-agent utility by 13.3% but cannot collapse the market. Mediation enables recovery even under sustained adversarial pressure. We define adversarial robustness as a mechanism's ability to sustain positive honest-agent utility under optimised attack, and find that Mediation is robust: it can be bent but not broken.