Skip to content
Preprint

Interaction Readiness: A Framework for Building and Evaluating AI Agents in Human Roles

Jul 2026 · 0 citations · 13 references
Computer Science

TL;DR

Interaction Readiness is introduced as a framework for specifying and evaluating that missing layer of performance in role-bearing AI agents, and its findings are translated into a specification template and audit procedures that product and engineering teams can apply before and after deployment.

Abstract

Product and engineering teams building role-bearing AI agents face an evaluation gap: an agent can produce accurate, safe, and fluent content while still failing the behavioral requirements of its assigned role. This paper introduces Interaction Readiness as a framework for specifying and evaluating that missing layer of performance. The framework separates content specifications, which govern what an agent knows and says, from interaction specifications, which define how an agent should conduct itself in a role-governed exchange. Interaction specifications require teams to define role purpose, authority boundaries, recurring situations, boundary cases, repair behaviors, and audit criteria before deployment. We operationalize interaction readiness through four agent operations: understanding purpose, calibrating authority, managing tone, and repairing breakdowns. Using StudyChat, a public dataset of student interactions with an AI tutoring agent, we show that content accuracy and interaction quality are independent dimensions: an agent may be factually correct while failing as a tutor, or interactionally sound while technically wrong. The most persistent failure is authority miscalibration: the agent often knows how to answer, but not whether, when, or how the tutor role permits it to answer. The paper translates these findings into a specification template and audit procedures that product and engineering teams can apply before and after deployment

View source

Similar papers

Review Aug 2026

Evaluating Skills, Not Just Agents: Agentic Continuous Evaluation of Skills

ACES (Agentic Continuous Evaluation of Skills), a repository-native framework for evaluating skills and product capability packages as executable agent artifacts, is presented.

Christopher Kevin, Narendran Raghavan, J. Puget et al. · 1 citation · ⚡1
Jul 2026

Authoring Agent Skills: A Software-Engineering Approach

This note argues that a skill is a software artefact and that its construction should follow software-engineering principles, with qualifications: single responsibility, separation of interface from implementation, low coupling, and economy in a shared token budget, together with behavioural evaluation in place of deterministic testing.

Giuseppe Destefanis · 1 citation
Review Aug 2026

AgentForge: An Immersive Role-Playing Platform for Learning Agentic Software Engineering

AgentForge is presented, an immersive learning system in which novices take on one of four software-engineering roles: Task Planner, Patch Author, Code Reviewer, or Test Runner, within a multi-agent code-repair workflow, which clarifies role-specific responsibilities, makes agent coordination and intermediate artifacts visible, and encourages novices to monitor and evaluate their decisions.

Zihan Fang, Yue-Ke Zhang, Yu Huang · 0 citations
#software testing Review Aug 2026

Model-Based Agentic Software Engineering

MAGE explains how externalized knowledge, bounded action, independent evaluation, and retained human authority can compose into a governed engineering environment, and proposes tests of when that environment turns commodity intelligence into durable engineering progress.

James C. Davis, Kelechi G. Kalu, Huiyun Peng et al. · 2 citations
Review Jul 2026

Specifying the Delegated-Autonomy Boundary: Requirements Engineering for Agentic AI

This paper proposes two complementary artifacts: an Agency Justification Record (AJR) helps teams decide when an agent is warranted over simpler alternatives and an Agentic Delegation Policy (ADP) captures what must be specified for safe and effective development.

Chetan Arora, Andreas Vogelsang, Abbishek Sharma · 0 citations
Preprint Aug 2026

Agent Gym: A Framework for Continuous Evaluation and Evolution of LLM Agents Through Human-in-the-Loop Feedback

Agent Gym is introduced, a modular, domain-agnostic framework that wraps any existing LLM-based agent in a continuous evaluation-and-evolution loop and introduces the Spec-to-Note Gap, an autoencoder-inspired view of agentic system transparency.

Pouya Ghiasnezhad Omran, Michael Zimmermann, Duncan Cambridge et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.