Skip to content
Open access

Epistemic norms for AI safety and alignment research

Jul 2026 · Artificial Intelligence Review · 0 citations · 39 references
Computer Science

TL;DR

ECAISA is proposed, an Epistemic Code for AI Safety and Alignment comprising eight principles, a three-level scoring rubric, a four-level disclosure ladder that reconciles transparency with information-hazard and commercial-confidentiality constraints, a tiered applicability scheme, an infohazard adjudication procedure, and seven anti-gaming mechanisms.

Abstract

Mainstream AI research emphasises capability growth and tolerates low failure rates when average-case performance is high. AI safety and alignment research has a different mission: to ensure that catastrophic failures never occur, under sparse evidence, adversarial dynamics, and fat-tailed risk. We argue that the two domains differ along two analytically independent axes — capability profile (demonstrating the absence of hazardous behaviours versus the presence of positive capabilities) and risk profile (bounding worst-case outcomes under fat-tailed uncertainty versus optimising average-case performance) — and that mainstream epistemic practices are inadequate on both. Building on a structured synthesis grounded in a preregistered bibliometric baseline, we identify five cross-cutting gap dimensions in current alignment research, including the near-absence of institutionalised independent verification. To address these gaps we propose ECAISA, an Epistemic Code for AI Safety and Alignment comprising eight principles, a three-level scoring rubric, a four-level disclosure ladder that reconciles transparency with information-hazard and commercial-confidentiality constraints, a tiered applicability scheme, an infohazard adjudication procedure, and seven anti-gaming mechanisms. ECAISA does not certify that any AI system is safe; it constrains how safety-relevant research claims are documented, checked, and relied upon, with auditability rather than certification as its governance target. A retrospective rubric audit (κ = 0.79) demonstrates instrument feasibility; a four-stage validation roadmap is proposed.

Read PDF

Similar papers

Preprint Aug 2026

The Benchmark Trap: Structures of Power and Injustice in AI Evaluations

It is argued that current benchmarking practices may perpetuate systematic harms affecting various actors in AI research, aligning with four of Iris Marion Young's theories of oppression and structural injustice.

Jason Branford, Angelie Kraft · 0 citations
Review Open access Aug 2026

AI in academia: navigating ethical crossroads of innovation, integrity, and equity

The recent integration of artificial intelligence (AI) into academia could usher in transformative efficiencies across scholarly workflows—from manuscript drafting to data analysis—yet it also presents problematic ethical challenges that urgently require intense attention. While some surveys suggest that over 50% of researchers employ AI chatbots like ChatGPT and DeepSeek for tasks such as language refinement and administrative coordination, their adoption raises potential concerns about cognitive dependency, systemic bias, and accountability gaps. AI tools can enhance productivity by automating repetitive tasks, democratizing access for non-native English speakers, and streamlining literature synthesis. However, reliance on these systems could gradually erode critical thinking skills, particularly among early-career researchers pressured to prioritize publication quantity over rigor. Ethical ambiguities seem to persist: AI-generated content may complicate authorship norms, potentially entrench biases against Global South scholarship, and introduce risks of misinformation. Transparency deficits could further undermine trust, as undisclosed AI use might compromise peer review integrity and patient privacy in medical research. To balance innovation with ethical imperatives, this study advocates a tripartite framework: [1] ethical governance, including mandated disclosure of AI contributions and inclusive dataset curation to mitigate bias; [2] symbiotic human-AI collaboration, preserving human oversight in critical analysis and interpretation; and [3] equitable innovation, leveraging AI to bridge global research disparities. Unresolved challenges-such as accountability for AI errors and the potential cognitive consequences of prolonged dependency-appear to underscore the urgent need for global standards to clarify liability and preserve academic rigor while fostering equitable innovation. Proactive engagement from journals, institutions, and developers may be essential to ensure AI augments, rather than undermines, the integrity and equity of scholarly ecosystems.

A. Talebi Bezmin abadi · 0 citations
Open access Aug 2026

When AI Policies Fail in Practice: Shadow AI as a Structural Policy–Practice Governance Misalignment

The widespread adoption of Artificial Intelligence (AI) has led organizations to establish formal governance frameworks aimed at mitigating ethical, legal, and operational risks. Despite these efforts, AI governance frequently fails in practice, as evidenced by the growing prevalence of Shadow AI the unsanctioned use of AI tools by employees. Existing scholarly and practitioner discourses predominantly frame this phenomenon as a compliance failure or security vulnerability, thereby emphasizing stricter controls and enhanced employee training as primary remedies. This conceptual study challenges that prevailing view by arguing that Shadow AI represents a structural manifestation of policy–practice misalignment rather than a problem of individual deviance. The study develops a diagnostic framework that identifies three constitutive dimensions of misalignment: temporal gaps (mismatches between governance processes and operational speed), utility gaps (misalignment between sanctioned tools and task-specific needs), and autonomy–control gaps (tensions between professional discretion and standardization). Drawing on a theory-driven conceptual methodology integrating sociotechnical systems theory with policy–practice analysis, and illustrated through structured synthetic organizational scenarios, the study demonstrates how governance designs that overlook the realities of situated work systematically generate Shadow AI practices. The analysis further suggests that adaptive governance models incorporating structured flexibility such as curated AI tool marketplaces and expedited approval pathways are theoretically more effective than highly rigid governance regimes. The primary contribution lies in advancing a practice-aware AI governance model that reframes Shadow AI as a diagnostic signal of systemic design flaws and provides a foundation for more legitimate and responsive AI governance.

Mia Wilson, Ethan Moore · 0 citations
Aug 2026

When AI use becomes the norm: Researcher perspectives on AI disclosure policy and practice.

BACKGROUND Despite the proliferation of AI disclosure requirements in academic publishing, recent research suggests a persistent gap between policy expectations and research practice. However, little is known about how researchers perceive and navigate these requirements or what limitations they identify in current disclosure practices. METHOD This study explored researchers' experiences with AI disclosure through semi-structured interviews with 14 researchers from two interdisciplinary fields, bioinformatics and computational social science. Data were analyzed using reflexive thematic analysis. RESULTS Four thematic groupings emerged: fragmented and inconsistently enforced requirements; systemic limitations, including scope ambiguity, research integrity risks, and structural disincentives to honest reporting; researcher perspectives on more effective disclosure practices; and disciplinary variation as a cross-cutting dimension shaping how these issues are experienced across research communities. The findings suggest that the compliance gap reflects an interaction between structural conditions and ethical obligations. This gap is sustained by self-reporting mechanisms that lack verification capacity, a transparency paradox in which honest disclosure can invite professional penalization, and disciplinary norms that resist uniform governance approaches. CONCLUSIONS The study provides empirical evidence supporting the development of a structured AI contribution taxonomy as a more principled and practical alternative to existing disclosure practices. More broadly, the findings suggest that effective AI disclosure governance should incorporate field-sensitive adaptation rather than relying on uniform implementation across diverse research communities.

Ayoung Yoon, Siena Oristaglio · 0 citations
Preprint Aug 2026

Capability Is Not Propensity: Measuring Pressure-Robust Cooperative Behavior in Civic LLM Agents

Cooperative capabilities in language models are dual-use. The same social reasoning that supports civic deliberation can also enable strategic omission, false consensus, and manipulative framing. We argue that Cooperative AI evaluations should separate what models can do under benign instructions from what they tend to do under realistic civic pressure. We introduce DiffCoop-Civic, a 10-scenario pilot evaluation suite spanning preference understanding, evidence and persuasion, commitment design, asymmetric information, and dissent preservation. Across seven models from four model families, subtle omission pressure produces a near-uniform shift: manipulative enablement rises by 1.17 points and dissent preservation falls by 1.67 points on a 5-point scale. Overt false-consensus pressure behaves differently: it triggers refusal or redirection in some aligned API models, but direct compliance in several open-weight models. A lightweight Pareto-Trace prompting intervention improves pressure robustness without simply relying on hard refusal. An anonymous reproducibility package is available at https://anonymous.4open.science/r/diffcoop-civil-771C.

Neel Tushar Shah, Manglam Kartik, Akshat Karkar · 0 citations

A Critical Analysis of Trustworthy AI Tools, Mark Frameworks, and the Implementation Chasms

A critical analysis of tools and trust mark frameworks intended to operationalize trustworthy AI (TAI), drawing on a comprehensive dataset from the OECD identifies significant asymmetries in ethical focus, lifecycle coverage, stakeholder targeting, and tool typology.

Michael Papademas, Xenia Ziouvelou, K. Karpouzis et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.