Skip to content
Conference

Understanding Engineering Challenges in AI Agent Frameworks for Software Development: An Empirical Study of GitHub Issues

Jul 2026 · Fall Joint Computer Conference · pp. 420-425 · 0 citations · 18 references

Abstract

AI agent systems increasingly support software engineering by extending large language models with capabilities such as planning, tool use, and coordinated execution, yet empirical evidence on the engineering challenges of building and maintaining such frameworks remains limited. To fill this gap, we conduct a large-scale empirical study of 3,864 closed GitHub issues from three representative repositories. We present a taxonomy of engineering challenges comprising 5 top-level categories and 21 subcategories, analyze the popularity and difficulty of these categories, and summarize 47 actionable solution strategies from resolved issue discussions and linked pull requests. These findings provide practical guidance for developers and framework providers, and offer an empirical basis for future research on AI agent engineering.

View source

Similar papers

Review Aug 2026

Developing LLM-based Multi-Agent Systems in Software Engineering: A Mixed-Method Experience Report

A comprehensive overview of the existing tools and frameworks for implementing MAS in software engineering and a set of lessons learned and challenges that can help researchers and practitioners to select a suitable MAS framework according to their needs are provided.

Maria Sâmyla Serafim de Oliveira, M. Ibiyo, Marco Gianrusso et al. · 0 citations
Preprint Aug 2026

An Exploratory Study of Agent Plans for Agentic AI Coding Tools in Open-Source Software

Overall, repository-preserved Agent Plans under these tool-specific directories appear to be a narrow but informative artifact for studying task intent and execution guidance in human-agent workflows.

M. Abubakar, Seyedmoein Mohsenimofidi, Jai Lal Lulla et al. · 1 citation
Book Open access Aug 2026

Agentic Software Engineering (SE 3.0): The Rise of AI Teammates

This workshop aims to define a roadmap for a world where AI Teammates and human developers build the future together, anchored by the launch of the AIDev dataset, which provides the empirical evidence needed to understand the behaviors of AI Teammates.

Hao Li, Haoxiang Zhang, Jie M. Zhang et al. · 1 citation
Review Open access Aug 2026

Knowledge Retrieval Architectures for AI-Assisted Software Development: From Static Context to Autonomous Development Agents

The integration of Large Language Models (LLMs) into software development workflows has fundamentally transformed how developers interact with knowledge repositories, codebases, and documentation. However, traditional knowledge retrieval mechanisms face significant challenges when applied to the dynamic, multi-modal, and contextually-rich environment of software engineering. This survey provides a comprehensive analysis of knowledge retrieval architectures specifically designed for AI-assisted software development, examining the evolution from static context provision to autonomous development agents. Building upon this systematic examination, the paper delineates critical research directions that advance the theoretical and practical foundations of AI-assisted software development.

Norbert Fijałek, I. Bluemke, Piotr Gawrysiak · 0 citations
Book Open access Jul 2026

Agents in the Wild: Where Research Meets Deployment

Through applied case studies in pharmaceutical discovery and financial systems, common design patterns that make agentic systems successful are analyzed, and practical mitigation strategies for failure modes are discussed, such as verification pipelines, fallback mechanisms, and human-in-the-loop supervision.

Grace Hui Yang, P. Venkit, Hooman Sedghamiz et al. · 0 citations
Review Aug 2026

Software Engineering for and with GUI Agent

GUI agents have advanced rapidly, producing a growing body of frameworks, benchmarks, and applications. However, this growth has outpaced the maturity of the field. GUI agents remain technically brittle, incompletely engineered, and insufficiently validated for sustained real-world use. They are evolving into closed-loop software systems. Within these systems, model reasoning is coupled with interface perception, execution feedback, recovery, and human oversight. This evolution calls for a software engineering perspective that remains largely absent from existing research. We address this gap by reviewing 336 GUI-agent papers from January 2018 to April 2026. Five research questions examine the research landscape, architectures, evaluation, software lifecycle concerns, and future opportunities. Our findings show that the field has expanded sharply since 2024, while mobile and web settings remain dominant. Architectures increasingly adopt modular perceive-reason-act loops, but recovery, human escalation, safety enforcement, and auditability remain underdeveloped. This architectural imbalance extends to evaluation. Evaluations are becoming more interactive, but they remain centered on task success and are difficult to compare across protocols. More broadly, existing studies provide limited support for testing beyond benchmarks and for maintaining agents after release. Observability, privacy engineering, and systematic human oversight are also underdeveloped. Together, these findings show that capability improvements alone cannot ensure deployment readiness. Future research should connect dependable execution with lifecycle-centered testing and reproducible evaluation. It should also integrate permission and privacy controls with cost-aware, human-centered governance. This integration is necessary to build dependable, maintainable, secure, and deployable GUI-agent systems.

Sheng-Cheng Yu, Yuchen Ling, Junyang Xing et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.