EvoMaestro: Toward Interpretable and Steerable LLM-Driven Program Evolution
Feng LiangSizhe ChengYikai LiRuijie HeXiaolin WenYong Wang
Oct 2026
Human-computer Interaction
Abstract
Recent program evolution systems use large language models (LLMs) to generate and iteratively improve populations of programs, producing striking advances in mathematics, algorithm design, and scientific computing. Yet these systems largely operate as fully automated black boxes. As populations grow, domain experts must make sense of the scores, code changes, reasoning, and algorithmic ideas across many programs, while current interfaces provide limited support for understanding or redirecting the evolution. We characterize this need to understand and steer populations of evolving algorithmic ideas as semantic oversight. A formative study with 8 domain experts yields six design requirements for this emerging human-computer interaction problem. We then propose a steerable program evolution framework that lets expert judgments shape subsequent evolution. Built on this framework, EvoMaestro is an interactive visual analytics system that organizes evolution information from population overview to source code. It helps experts locate and compare noteworthy programs, guide evolution with natural language, combine promising ideas, and prune unproductive directions. A seven-day system demonstration illustrates how an expert applied these capabilities throughout a long-running evolution process. A within-subjects study with 12 participants shows that EvoMaestro improves users' understanding of evolution processes, reduces cognitive workload, and supports expert steering. These findings suggest that preserving expert agency in open-ended LLM-driven search requires support for both informed judgment and the ability to shape subsequent automation.
This publication proposes a definition and a classification of agile software development approaches and analyses ten software development methods that can be characterized as being "agile" against the defined criterion.
P. Abrahamsson, O. Salo, Jussi Ronkainen et al.· arXiv.org· 727 citations· ⚡54
The study shows that agile practices improve both informal and formal communication, but indicates that, in larger development situations involving multiple external stakeholders, a mismatch of adequate communication mechanisms can sometimes even hinder the communication.
M. Pikkarainen, Jukka Haikara, O. Salo et al.· Empirical Software Engineeri...· 401 citations· ⚡48
The results indicate that software engineering work practices are chosen opportunistically, adapted and configured to provide value under the constrains imposed by the startup context.
Nicolò Paternoster, Carmine Giardino, M. Unterkalmsteiner et al.· Information and Software Tec...· 394 citations· ⚡54
The perception of the impact of agile methods is predominantly positive, and several challenge areas were discovered, but based on this study, agile methods are here to stay.
M. Laanti, O. Salo, P. Abrahamsson· Information and Software Tec...· 260 citations· ⚡20
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduOct 7, 2026
Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.
Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity. The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.
MIT News · Artificial Intelligence· news.mit.eduSep 30, 2026
Able to defeat top-ranked human players and more efficient than other models, the new system could help decision-makers in military maneuvers or business negotiations.