It is argued that AI systems used in conducting foreign policy tasks - broadly enacting Tatecraft'- should be a priority test case for technical AI governance research, and an ECOSYSTEM review highlights the asymmetric focus on ASSESSMENT features over ACCESS, VERIFICATION, SECURITY, and OPERATIONALIZATION.
Abstract
We argue that AI systems used in conducting foreign policy tasks - broadly enacting'statecraft'- should be a priority test case for technical AI governance research. In enacting foreign policy, we refer to the formulation and implementation of external objectives by political actors. Statecraft is a high-consequence deployment domain, with extreme downside risks and structural properties that standard evaluation practices handle poorly. These features include partial observability, unbounded action spaces, contested ground truth, and multidimensional objectives. This paper advocates for a literature-grounded research agenda. Our contribution is threefold: (i) a claim about the structural conditions of foreign policy that combine catastrophic tail risk with technical evaluation complexities, (ii) an ECOSYSTEM review that highlights the asymmetric focus on ASSESSMENT features over ACCESS, VERIFICATION, SECURITY, and OPERATIONALIZATION, and (iii) a demand-side evaluation framework that decomposes foreign-policy workflows into bounded, evaluable sub-tasks with human recombination. As AI systems are already being deployed in the conduct of war and peace, amid limited public evaluation infrastructure from the technical AI governance community, this agenda is an urgent priority.
Current AI risk assessment methods in industry rely heavily on subjective judgment. Yet recent U.S. AI policy frameworks, including Executive Order 14319 and America's AI Action Plan, mandate that AI systems be "free from ideological bias" and pursue "objective truth". While NIST provides systematic risk evaluation guidance, current AI governance frameworks are not designed to meet such objectivity requirements, instead offering flexibility that accommodates implementation across various contexts. Organizational AI risk practices rely heavily on subjective likelihood and impact scoring, compounded by subjectivity introduced when results are translated from technical teams to executives. Literature addressing this gap spans three disconnected streams: (1) AI governance frameworks relying on subjective risk assessment, (2) technical ML bias measurement focused on algorithmic fairness, and (3) organizational implementation approaches failing to translate governance principles into operational practice. Previous approaches lack links between risk analysis, objectivity requirements, and actionable guidelines, while policy frameworks remain too broad for practical application. Missing is research systematically bridging policy objectivity requirements with sociotechnical challenges of organizational risk assessment. To address this gap, we propose a framework transforming observable organizational factors into measurable risk indicators. Drawing from socio-technical systems theory and established risk taxonomies, we decompose traditional likelihood and impact metrics into specific, observable criteria replacing subjective estimates vulnerable to biases, addressing organizational tendencies to underestimate risks.
Unknown authors· Proceedings of IASEAI Confer...· 0 citations
This paper introduces a policy analysis framework for systematic, transparent assessment of AI governance proposals in an evolving and contested regulatory landscape. AI policy debates often collapse into binary positions that obscure underlying tradeoffs and normative assumptions. The framework structures policy analysis around multiple policy attributes, allowing users to surface priorities and tensions without prescribing outcomes. We use a mixed-methods approach that integrates qualitative insights from subject matter experts with computational text analysis to inform the design of policy attribute rubrics. This quantifies the relative emphasis of different policy objectives and presents them through comparative visualizations that support interpretability and cross-policy comparison. The paper also examines the use of commercial LLMs for rubric-based policy analysis, benchmarking their outputs against a domain-trained rubric-calibrated model with explicitly defined analytical assumptions. Rather than assessing policy effectiveness or desirability, the framework focuses on relevance and alignment across attributes. By making analytical assumptions explicit, including attribute selection, rubric construction, and weighting schemes, the framework enables users to evaluate whether its embedded priorities align with the users'own normative commitments. The approach is jurisdiction-agnostic and intended to support policymakers, analysts, and researchers navigating complex AI governance environments. Contributions: (1) multidimensional policy assessment through empirically grounded rubrics that surface tradeoffs rather than resolving them; (2) a transparent hybrid methodology combining feedback from subject-matter experts with computational validation; and (3) use of domain-trained rubric-calibrated models as a benchmark for comparing different general-purpose large language models.
Paulo Carvao, C. M. Verdun, Isabel L Adler et al.· arXiv.org· 0 citations
The widespread adoption of Artificial Intelligence (AI) has led organizations to establish formal governance frameworks aimed at mitigating ethical, legal, and operational risks. Despite these efforts, AI governance frequently fails in practice, as evidenced by the growing prevalence of Shadow AI the unsanctioned use of AI tools by employees. Existing scholarly and practitioner discourses predominantly frame this phenomenon as a compliance failure or security vulnerability, thereby emphasizing stricter controls and enhanced employee training as primary remedies. This conceptual study challenges that prevailing view by arguing that Shadow AI represents a structural manifestation of policy–practice misalignment rather than a problem of individual deviance. The study develops a diagnostic framework that identifies three constitutive dimensions of misalignment: temporal gaps (mismatches between governance processes and operational speed), utility gaps (misalignment between sanctioned tools and task-specific needs), and autonomy–control gaps (tensions between professional discretion and standardization). Drawing on a theory-driven conceptual methodology integrating sociotechnical systems theory with policy–practice analysis, and illustrated through structured synthetic organizational scenarios, the study demonstrates how governance designs that overlook the realities of situated work systematically generate Shadow AI practices. The analysis further suggests that adaptive governance models incorporating structured flexibility such as curated AI tool marketplaces and expedited approval pathways are theoretically more effective than highly rigid governance regimes. The primary contribution lies in advancing a practice-aware AI governance model that reframes Shadow AI as a diagnostic signal of systemic design flaws and provides a foundation for more legitimate and responsive AI governance.
Mia Wilson, Ethan Moore· Journal of Management and In...· 0 citations
It is argued that current benchmarking practices may perpetuate systematic harms affecting various actors in AI research, aligning with four of Iris Marion Young's theories of oppression and structural injustice.
Public services face growing pressure to adopt artificial intelligence (AI) to close the gap between rising demand and falling resources. That pressure has intensified with general-purpose AI (GPAI): AI built on large language models that can be directed by prompt alone to perform an effectively unbounded range of tasks. We argue that the properties that make these models attractive - their generality, accessibility, and low deployment cost - undermine the conditions under which AI safety has historically been pursued. The safety concepts that public service governance frameworks foreground - accuracy, bias, explainability, and accountability - were made tractable by narrow, purpose-built AI, and the mitigations that guidance documents prescribe presuppose exactly what GPAI removes. Accuracy cannot be quantified over unbounded outputs. Bias cannot be disaggregated when outputs are free-text judgements rather than categorical predictions. Explainability gives way to the appearance of explanation, and accountability erodes as outputs are optimized to persuade. We develop this through the case of policing, where the consequences of governance failure are most severe, and show why the same failure is likely to recur across other public services. The two mitigations that dominate policing AI strategy - expert evaluation and human-in-the-loop oversight - both rest on assumptions that GPAI violates. Safety assurance thus shifts from an intrinsic feature of building an AI tool to an optional add-on. We recommend a clear taxonomic distinction between narrow and general-purpose AI in governance documentation, a preference for technological parsimony, a pause on operational deployment of GPAI in policing until adequate evidence exists, and a coordinated national safety infrastructure with the authority to generate that evidence and determine when responsible deployment is achievable.
When the European Union (EU) adopted its AI Act (AIA) in 2024, hopes were high that EU rules would diffuse globally through a “Brussels Effect.” We investigate how structural and contextual factors in the field of AI governance affect the likelihood and transformative potential of a Brussels Effect there. Because the AIA is still too young to allow a straightforward ex post analysis, we combine two inferential strategies: we compare AI as a governance challenge to other cases of successful EU rule export, identifying conditions that promote or obstruct global EU rule export. And we empirically canvas the rule-setting dynamics that have emerged since the AIA’s first legislative draft in 2021. Against the initially optimistic tenor of Brussels policy discourse from the time of the AIA’s adoption, we advance four reasons to be skeptical of big EU influence: first, the EU’s regulatory capacity is constrained by informational disadvantages, reliance on external standard-setters, and enforcement delays. Second, competitive pressures have led to pre-emptive concessions, watering down EU rules before they could travel. Third, the relative vagueness of EU rules means that de factor harmonization of rules and practices across borders remains shallow at best. Fourth, rule diffusion in AI is above all meaningful when it reaches countries that themselves are leading AI powers, notably the United States and China, and it is precisely there that EU influence seems negligible so far.
Unknown authors· Digital Society· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.