Skip to content
Review

AgentForge: An Immersive Role-Playing Platform for Learning Agentic Software Engineering

Aug 2026 · 0 citations · 41 references
Computer Science

TL;DR

AgentForge is presented, an immersive learning system in which novices take on one of four software-engineering roles: Task Planner, Patch Author, Code Reviewer, or Test Runner, within a multi-agent code-repair workflow, which clarifies role-specific responsibilities, makes agent coordination and intermediate artifacts visible, and encourages novices to monitor and evaluate their decisions.

Abstract

Agentic AI is increasingly used to coordinate planning, implementation, review, and testing in software development, yet it often offers limited transparency into its decisions and interactions. Many such systems also assume that users can effectively guide the AI's decisions and validate its outputs. This assumption poses a particular challenge for novices, who must simultaneously learn how agentic AI works, how to collaborate with it effectively, and how to evaluate its outputs critically. To address this challenge, we present \textit{AgentForge}, an immersive learning system in which novices take on one of four software-engineering roles: Task Planner, Patch Author, Code Reviewer, or Test Runner, within a multi-agent code-repair workflow. In each practice session, the novices perform their chosen role while AI agents perform the remaining three. Through role-based scaffolding and metacognitive support, AgentForge clarifies role-specific responsibilities, makes agent coordination and intermediate artifacts visible, and encourages novices to monitor and evaluate their decisions. In a study with 37 novice developers, participants achieved high task-completion rates with AI-agent support. However, interaction demands differed significantly across practices: the Code Reviewer practice required more interaction turns, reroutes, and completion time ($p_{\mathrm{adj}} = .004$) and was perceived as the most challenging. Participants nevertheless reported significant gains in their understanding of software repair and agent collaboration ($p_{\mathrm{adj}}<.001$). These findings suggest that AgentForge can help novices develop practical software-engineering skills while learning to collaborate with agentic AI more critically and effectively.

View source

Similar papers

A Goal-Oriented Agentic Framework For Collaborative Branching Human-Robot Interactions

It is suggested that the benefit of the agentic framework lies primarily in interaction quality rather than conversational efficiency, and the agentic architecture displayed robustness by recovering from non-normative inputs while maintaining strict goal alignment.

Morten Roed Frederiksen · 1 citation
Jul 2026

Authoring Agent Skills: A Software-Engineering Approach

This note argues that a skill is a software artefact and that its construction should follow software-engineering principles, with qualifications: single responsibility, separation of interface from implementation, low coupling, and economy in a shared token budget, together with behavioural evaluation in place of deterministic testing.

Giuseppe Destefanis · 1 citation
Preprint Aug 2026

ETA: A New Agentic Paradigm for Embodied Tasks

The Embodied Task Agent is introduced, a new paradigm for extending digital agents into the physical world, and OpenETA is released as its open-source implementation, which provides replaceable Planners, composable Tools and Skills, auditable memory, replayable trajectories, and common interfaces for simulation and real robots.

Yitong Chen, Zezheng Huai, Sixian Li et al. · 1 citation
Review Open access Sep 2026

AI Agents Can Now Navigate and Complete LMS Tasks: A Call for Pedagogical Innovation

Autonomous artificial intelligence (AI) agents can now log into a learning management system, read course materials, and complete unproctored, asynchronous assessed work end-to-end with no student involvement. We document that capability and trace its consequences for assessment validity. Three demonstrations on a live undergraduate course supply the evidence: two quiz completions, one in approximately 12 minutes, one in under 5, and a third in which the agent fabricated credible personal reflection for a discussion board. The wider public record includes at least 15 documented agent runs across three platforms and seven tools. We apply Kane’s argument-based validity framework: agent completion removes the attribution on which every inference in Kane’s chain depends. Everything downstream, from course grades to the evidence chains behind program review and accreditation, rests on support that is no longer there. The failure concerns validity rather than integrity: an institution can punish misconduct and still lack grounds for the scores it reports. Collective accreditor guidance addresses institutional uses of AI in evaluation and does not yet reach the agentic case. Audience data from the underlying conference session show attendees already recognizing both the vulnerability and the gap in institutional guidance. Polled attendees most often named online quizzes as agent-completable, with discussion-based work close behind. Majorities in both listings were working without settled written guidance. The response defended here is design rather than detection: four principles for verified human presence, low-effort changes faculty can adopt now, and the assurance levers assessment professionals already operate.

Stavros P. Hadjisolomou, R. El-Haddad · 0 citations
Open access Aug 2026

Tool-Augmented Language Agents with Iterative Self-Critique for Complex Task Planning

The research details a comprehensive methodological framework, formalizing the probabilistic decision-making and critique generation processes and indicates that integrating reflective cognition paradigms with modular toolsets is essential for deploying autonomous language agents in high-stakes, real-world applications.

Mabel Kwok · 0 citations
Review Aug 2026

Evaluating Skills, Not Just Agents: Agentic Continuous Evaluation of Skills

ACES (Agentic Continuous Evaluation of Skills), a repository-native framework for evaluating skills and product capability packages as executable agent artifacts, is presented.

Christopher Kevin, Narendran Raghavan, J. Puget et al. · 1 citation · ⚡1

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.