Skip to content
Open access

Wild, Thick, and Wicked: Situated Evidence on AI-In-Use for Decisions About Deploying AI Systems

Sep 2026 · Social science computer review · 0 citations · 66 references

TL;DR

A real-world AI evaluation framework focused on AI-in-use: how people actually appropriate, adapt, and work around AI systems in context, and what consequences follow over time is proposed.

Abstract

Organizations are adopting generative AI faster than the evidence base needed to govern it. Existing evaluation tools such as benchmarks, alignment scores, and safety tests were built for model development, not for judging whether systems will create value, introduce friction, or shift risk in specific real-world settings. As a result, there is little systematic evidence about how AI behaves once it is embedded in everyday work. This paper proposes a real-world AI evaluation framework focused on AI-in-use: how people actually appropriate, adapt, and work around AI systems in context, and what consequences follow over time. Instead of treating variability across users, tasks, and settings as noise to be controlled away, the framework treats that variation as the central source of deployment-relevant evidence. It sets out four design principles for producing decision-ready evidence at scale and proposes a shared evaluation architecture combining a structured observation environment, a metrics hub, and reusable consortium models that summarize system behavior across contexts. Rather than replacing traditional benchmarks, this framework adds a sociotechnical evidence layer that connects model capabilities to the organizational and practitioner level outcomes where deployment decisions are actually made.

Read PDF

Similar papers

#explainable ai Review Open access Sep 2026

Human-In-The-Loop Decision Systems: Advances and Future Directions

It is argued that treating the human and the model as a single joint cognitive system is the central design principle for the next generation of decision systems.

Ashore-Onisemo Funmilayo · 0 citations
Open access Sep 2026

Borrowed or kept? How self-generation-first and direct-adoption AI workflows shape immediate unaided reasoning

A three-arm experiment that varies the manner of AI use against a no-AI anchor and measures performance 15-20 min later, once the tool has been removed, results in an immediate near-transfer decrement rather than demonstrated lasting de-skilling.

Cheng-Hai Liu, Jia-Yao Guo, Rong-Xia Gao et al. · 0 citations
Open access Sep 2026

Balancing Trust and Deliberation in Human–AI Decision Support: The Effects of Explainable AI and Cognitive Forcing Functions

Reliance on AI systems for decision support is expanding into domains where mistakes carry real consequences, which makes it important that users can weigh AI suggestions against their own judgment rather than deferring to them by default. Prior work has found that human-AI teams sometimes underperform AI alone, a patt...

Oliver Henderson · 0 citations
Open access Aug 2026

Do you know what your AI agent can do on its own?

A two-dimensional design space is introduced in which both dimensions are organised into five operational levels, making the coupling explicit and navigable, and six architectural tactics for adjusting a deployment’s position within it are proposed, offering a shared vocabulary for compliance-aware agentic AI design.

D. Safin, Dian Baltaa, Timon Sengewaldb et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Rethinking Data Quality for AI-Driven Systems: Evidence from Practitioner Interviews

Data quality research has usually treated data as an input that is stored, processed, and validated. In AI-driven software-intensive systems, data also shapes model behavior, evaluation, and lawful use. Empirical evidence remains limited on how practitioners define, assess, and manage quality under these conditions. We...

Hariharan Gopinath, Jan Bosch, H. Olsson · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.