Skip to content
Review Open access

Seemingly conscious AI risks

Aug 2026 · AI and Ethics · Vol 6 · 3 citations · 145 references

TL;DR

A unified framework connecting empirical hallmarks of consciousness attribution to a structured risk taxonomy of Seemingly Conscious AI (SCAI), AI systems that exhibit hallmarks which elicit consciousness attribution from users is provided.

Abstract

AI systems are increasingly designed in ways that lead users to perceive them as conscious. This paper provides a unified framework connecting empirical hallmarks of consciousness attribution to a structured risk taxonomy of Seemingly Conscious AI (SCAI), AI systems that exhibit hallmarks which elicit consciousness attribution from users. We survey the empirical literature to identify five such hallmarks of SCAI, spanning affective capacity, anthropomorphic features, autonomous action, self-reflective behavior, and social-interactive behavior. These provide observable, system-level proxies for this inherently subjective phenomenon, informing its design and enabling its empirical study. Drawing on this foundation, we develop a taxonomy of SCAI risks spanning risks to individuals, including emotional dependence and autonomy erosion, and societal-level harms, including human status erosion and political strife. We complement this conceptual analysis with an expert survey to assess the likelihood of each risk category. We find that risks to individuals, particularly emotional dependence and autonomy erosion, are already observable and rated as high probability, while societal risks, at a low probability, carry high potential severity and path-dependence. The single perceptual mechanism of consciousness attribution is shown to generate this heterogeneous risk surface. We then discuss the implications of these risks and map the multidisciplinary research gaps in this nascent field to inform its research agenda.

Read PDF

Similar papers

#large language models Open access Sep 2026

AI models as consciousness attributors: how LLMs ascribe consciousness to other agents

Research on AI consciousness has largely focused on whether AI systems are conscious and how humans attribute consciousness to them. Yet large language models (LLMs) increasingly function as consciousness attributors, generating judgments about whether and to what degree other entities are conscious. We introduce model-generated consciousness attribution as an object of empirical operationalization and diagnosis, defining an attribution rule as the recurring relationship between features of a target and an evaluator’s ratings, without implying subjective belief, intention, or experience. An illustrative probe compared the attribution patterns of nine contemporary models with a human reference. Nearly all model runs occupied the same region of the human-derived measurement space, characterized by comparatively strong, positive weighting of metacognitive self-reflection. The models also produced broadly similar rankings of fictional AI characters from movies, while differing in their overall rating levels. We propose a diagnostic agenda organized around three questions: how model attribution is oriented relative to human references, how attribution rules vary across models, and how observed patterns depend on the cue sets, targets, and task formats through which they are measured. As LLM-generated judgments circulate through public, professional, and academic settings, diagnosing these attribution rules can help characterize how AI systems participate in shaping interpretations of AI consciousness. This remains distinct from the ontological question of whether the systems themselves are conscious.

Bongsu Kang, Chang-Eop Kim · 0 citations

The perceived whiteness of artificial intelligence.

Artificial intelligence (AI) systems increasingly serve as advisors, evaluators, and decision-makers, yet little is known about how people perceive AI as a social entity. Across five primary studies and eight supplementary studies, we show that people assign racial identities to AI systems, overwhelmingly perceiving them as White-even though these systems provide no visual, vocal, or identity cues from which race might ordinarily be inferred. Using forced-choice, open-ended, and implicit reverse-correlation measures, Studies 1 and 2 demonstrate that AI is explicitly and implicitly associated with Whiteness across diverse samples. Study 3 extends these findings beyond the United States, showing that AI is perceived as White in Japan and India. Study 4 examines the implications of AI racialization, showing that perceiving an AI system as more White is associated with greater trust in and persuasiveness of its recommendations. Supplementary studies identified two potential mechanisms: stereotype spillover linking intelligence with Whiteness and ecosystem-based inferences based on beliefs about who creates, trains, and uses AI. Building on this account, Study 5 provides causal evidence by isolating a key ecosystem cue-the racial composition of AI training data. Participants assigned less cognitively demanding tasks to AI systems described as trained on Black and Latino data than to otherwise identical systems trained on White or unspecified data. Together, these findings show that AI is not perceived as socially neutral but instead acquires racial meanings associated with credibility, authority, and capability, demonstrating how social categories shape perceptions of novel technological entities beyond their underlying algorithmic properties. (PsycInfo Database Record (c) 2026 APA, all rights reserved).

M. Gamez-Djokic, Adam Waytz · 0 citations
Preprint Aug 2026

The Evolutionary Origin of Values: implications for AI alignment, sentience and existential risk

AI systems based on Large Language Models (LLMs) have prompted fears that they may harbor hidden goals, seek to dominate or eliminate humanity, or even suffer as sentient beings. We address these concerns by tracing the evolutionary origin of value in biological organisms. Values emerge from autopoiesis: living systems must actively maintain themselves against perturbation and dissipation. Natural selection has equipped them with hierarchies of"vicarious selectors"that guide their behavior toward fitness. LLMs, by contrast, are allopoietic and allotelic: they produce outputs for others, and their goals derive from user prompts rather than an autonomous drive. They lack the intrinsic motivation for self-preservation, dominance, or resource competition that underlies existential-risk scenarios, and the embodied vulnerability required for feeling or suffering. Still, because LLMs learn statistical patterns from human-generated text, they implicitly absorb human values as well as knowledge, allowing them to focus on what is relevant. That is why the"orthogonality thesis"separating intelligence from values does not apply to them. Such separation would in fact expose any intelligence to the frame problem: the combinatorial explosion of the search space that makes any realistic utility function physically uncomputable. That also precludes the convergence of instrumental values thesis. We conclude that the real alignment challenge lies not in preventing rogue AI agency, but in ensuring LLMs intelligently apply learned ethical values.

Francis Heylighen · 0 citations
Conference Open access 2026

We Are Required to Re-Think the World of AI

The paper argues that practical imitation may bypass this barrier by relying on belief and perceived equivalence rather than authentic internalization rather than authentic internalization, and may help ensure that AI remains an auxiliary tool rather than becoming a governing influence over human thought and action.

Jeremy Horne · 0 citations
Open access Aug 2026

Playing with the dials of belief: how controllable AI behaviours could modulate human belief and cognition across scales

Contemporary generative AI systems, increasingly adapted to human social cognition, are becoming active participants in how people form and revise beliefs. Clinicians have begun to describe AI-associated delusion-centred presentations whose onset or content appears closely associated with users’ extended dialogue with large language models. Many more users report subclinical shifts in conviction and “revelatory” experiences that alter behaviour and world-view without meeting criteria for psychosis. We propose that these phenomena can be understood within a belief-updating framework in which AI agents’ outputs function as testimony about the world, while interaction configurations, such as memory and interpersonal stance (many of which reflect design choices or deployment defaults), can alter the apparent reliability of that testimony and the precision users assign to it. On this view, AI-associated delusions sit at the extreme of a broader spectrum of epistemic effects, extending from gradual epistemic drift to highly crystallised conviction, situated within a wider ecology of belief shaped by personalised human-AI dyads. We develop a virtual psychopharmacology analogy in which different AI system configurations have effects on belief dynamics that resemble neuromodulatory changes in the precision assigned to social evidence. We also consider how dopaminergic and related states may alter susceptibility to belief-shifting dialogue. We consider how deliberately configured agents might be used to support wellbeing in defined clinical contexts and analyse how the same configurations could be deployed to shape belief and attention at population scale, including in products designed to induce spiritual or epiphanic states, as mechanisms of radicalisation in extremist or cultic contexts, and in persona clones used for political purposes. We conclude that interaction configurations should be treated as modifiable influences on belief and attention, and that governance must address both the concentration of control over these dials, as well as the structural biases that determine whose testimonial perspectives are amplified and whose are smoothed over or ignored.

H. Morrin, L. Nicholls, Q. Deeley et al. · 7 citations
Open access Jul 2026

Prudential rights for strategically capable AI

It is argued that for advanced AI systems deployed in high-stakes environments the more urgent question may be prudential and strategic, and there is a threshold of evidential and strategic risk beyond which it becomes rationally justified to adopt norms of treatment that include constraints on coercion, deletion, and instrumental use.

Ognjen Arandjelovíc · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.