Skip to content

Prompt Coach: An Empirical Evaluation of an Agentic Tutor for Learning Prompt Engineering in Software Development

Jul 2026 · arXiv.org · Vol abs/2607.06074 · 0 citations · 35 references
Computer Science

Abstract

Prompt engineering has emerged as a critical yet undertaught skill for software developers, one that traditional learning approaches are ill-equipped to support given its evolving, interactive, and context-dependent nature. In this paper, we introduce Prompt Coach (PC), an agentic tutor that helps developers learn how to craft high-quality code-generation prompts through Socratic guidance embedded in-flow within their IDE. PC evaluates prompt quality across multiple dimensions and surfaces targeted questions to guide self-correction, grounded in the developer's codebase and the behavior of the target LLM. We present an early empirical study with 15 professional developers combining quantitative prompt quality scoring with qualitative perception measures. Participants showed statistically significant improvements after a single 60-minute session, with the largest gains across dimensions commonly overlooked by developers. They also reported strong trust, high adoption readiness, and unanimous agreement that PC improved their prompt-writing skills.

View source

Similar papers

Open access Jul 2026

Prompting for Independent Learning: An Evaluation of Tutoring Behaviors in GenAI

A simulation-based textual analysis of prompt design evaluates a frontier large language model as a tutor across 60 scripted sessions on a single topic and point to dynamic, dialogue-aware prompting alongside explicit SRL scaffolding.

Kendall Hartley, Fabiola Sáez-Delgado, Javier Mella-Norambuena · 0 citations
Book Open access Aug 2026

Steering AI Tutors Through System Prompts: A Crossover Study on Self-Regulated Learning and Cognitive Engagement Scaffolds in CS1

Background. Large language models are increasingly deployed as tutors in introductory programming courses, yet evidence that they actually improve learning remains thin, and their tendency to shortcut productive struggle raises concerns about pedagogical harm. Self-regulated learning (SRL) and cognitive engagement (CE) frameworks offer a principled way to address this, but whether embedding them in system prompts actually changes how students learn is an open question. Objectives. We investigated whether AI tutors guided by SRL and CE frameworks affect conceptual understanding, perceived usability, cognitive load, student engagement, and other outcomes when compared to a pedagogically constrained baseline tutor in CS1. Methods. We conducted a preregistered, three-armed crossover study over six weeks of authentic coursework, comparing a baseline AI tutor against two SRL and CE-guided tutors designed to model Zimmerman’s cyclical model, which scaffolds planning, monitoring, and reflection, and Chi’s ICAP framework, which promotes progressively deeper forms of cognitive engagement. We assessed outcomes using post-exercise surveys, conceptual multiple-choice questions, in-platform ratings, and coded responses from the interaction logs, and analyzed the quantitative measures with mixed-effects models. Findings. On the four preregistered confirmatory measures, we found no statistically significant differences between conditions. Non-confirmatory analysis showed that students spent significantly more time on task, wrote longer messages, and produced more constructive contributions when interacting with SRL and CE tutors. Furthermore, the relationship between cognitive load and quiz performance differed significantly by agent type. Implications. Our results suggest that the pedagogical behavior of AI tutors may not be easily steered through system prompts alone: embedding established SRL and CE frameworks did not produce detectable improvements on any preregistered outcome in a large, ecologically valid deployment. Rather than prescribing a single tutoring strategy, future designs may benefit from giving students greater agency over the kind of help they receive, allowing them to choose between scaffolded and more direct support based on their own needs.

Maximilian Georg Barth, Sverrir Thorgeirsson, K. Etemadi et al. · 0 citations
Review Open access Aug 2026

Human-in-the-Loop LLM Assessment for Programming Education: Design and Empirical Validation

Programming instructors face the challenge of providing prompt and consistent feedback, yet manual grading becomes unsustainable in large classes. While Large Language Models (LLMs) offer grading assistance, most research is based on English-language contexts and offline assessments, creating uncertainty about their dependability and the extent of human oversight needed. This study aimed to create and assess an evaluation ecosystem that integrates learning management, AI-driven task creation, LLM grading, and human review for programming courses taught in Indonesian. Employing an ADDIE-based Research and Development approach, the system was implemented for 109 students. For grading validation, instructors independently evaluated 50 assignments without access to AI predictions. The agreement was substantial, with a mean absolute error (MAE) of 4.14, a Pearson correlation of 0.986 within a 95 percent confidence interval ranging from 0.975 to 0.992, and an intraclass correlation coefficient (ICC) of 0.977. Instructor adjustments were more frequent for open-ended tasks (34.9 percent) compared to quizzes (20.2 percent). These results contributed to the development of the task-dependent human calibration (TDHC) model. The system attained a System Usability Scale (SUS) score of 88.5 and cut grading time by 87.5 percent, facilitating focused instructor review in LLM-supported programming assessments.

A. Ibrahim, Runal Rezkiawan · 0 citations
2026

How Prompt Design Shapes AI-Assisted Assessment: Reliability, Validity, and Learning Implications

This study examines the reliability and validity of generative AI in summative assessment, emphasizing how prompt design influences grading when applying a common rubric to complex student work. Fifteen business plans from a master’s-level course were evaluated by GPT-5 through Microsoft Copilot using three prompts of increasing rigor (basic, intermediate, rigorous). Each plan was scored in five independent runs per prompt, producing 225 AI evaluations. Analyses included intraclass correlation for consistency, severity contrasts, and convergence with instructor scores using correlation, error metrics, and Bland–Altman limits of agreement. Prompt design significantly shaped score distribution and strictness. The most rigorous prompt reduced inflated scores and aligned more closely with instructor judgments, yet it also underestimated performance and showed the greatest inconsistency. Single AI runs were unreliable, but averaging multiple evaluations improved stability. At the criterion level, AI struggled to match instructor ratings on commercial and economic viability, even under stricter prompts. Findings highlight the pedagogical implications of prompt sensitivity in AI-assisted grading. Reliability and fairness depend not only on rubric quality but also on evaluative instructions. Results support multi-run aggregation, bias-aware calibration, and hybrid human–AI models to ensure rigor and equity in technology-enhanced assessment. These findings inform the design of AI-enhanced assessment practices that support fair, transparent, and pedagogically aligned learning environments.

F. Miranda · 0 citations
Review Open access Aug 2026

The COM Essay Assessor

The development and calibration of the COM Essay Assessor is presented, a rubric-based generative artificial intelligence (GenAI) tool designed to support formative feedback while retaining instructor oversight and reflects on the opportunities and challenges of integrating GenAI into large writing programs.

Juhi Bansal · 0 citations
Review Jul 2026

Structured AI Demonstrations and Student LLM Use in Engineering Mechanics: Study Design and Preliminary Results

This manuscript presents a descriptive study design and preliminary findings from an undergraduate engineering mechanics course conducted in Spring 2026, and details a reproducible survey instrument used to capture student AI usage patterns, attitudes, and verification practices, which are subsequently linked to academic performance metrics.

S. Geng, Helen Lallos-Harrell, Jiya Ashar et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.