Apr 2026· AGENT@ICSE· pp. 152-156· 0 citations· 20 references
Computer Science
TL;DR
Preliminary results indicate performance differences across LLMs, suggesting that model choice influences coverage, consistency, and hallucination rates.
Abstract
This paper presents MARARE, a real-time multi-agent system that transforms meeting dialogues into structured software requirements. One agent interacts with participants, while background agents extract and verify requirements collaboratively. Evaluation using the LLM-as-a-Judge method across five meetings (5–8 minutes each) shows a mean coverage of 80.0 ± 11.2 % (mean ± SD), semantic similarity of 0.86 ± 0.05, and hallucination rate of 14.3 ± 6.2 %. Preliminary results indicate performance differences across LLMs, suggesting that model choice influences coverage, consistency, and hallucination rates.
This workshop aims to define a roadmap for a world where AI Teammates and human developers build the future together, anchored by the launch of the AIDev dataset, which provides the empirical evidence needed to understand the behaviors of AI Teammates.
Hao Li, Haoxiang Zhang, Jie M. Zhang et al.· Proceedings of the 32nd ACM...· 1 citation
AgentForge is presented, an immersive learning system in which novices take on one of four software-engineering roles: Task Planner, Patch Author, Code Reviewer, or Test Runner, within a multi-agent code-repair workflow, which clarifies role-specific responsibilities, makes agent coordination and intermediate artifacts visible, and encourages novices to monitor and evaluate their decisions.
An AI Task Agent system based on a Telegram chatbot, integrated with n8n as a workflow automation platform and Supabase as the database is developed that can improve the effectiveness of students’ task management through more practical, responsive, and structured interactions.
Huzain Azis, N. Widiyanti· Indonesian Journal of Data a...· 0 citations
As AI agents become integral to business workflows, establishing guiding user experience (UX) principles is crucial for ensuring user trust and successful adoption. To address this, our study uses a multi-method approach - combining participatory design workshop, paper-and-pencil, expert review, meta-analysis, and in-depth interviews - to identify and validate a design framework of eight core UX principles for human-AI agent interaction in the workplace. Together with their underlying criteria, these principles provide actionable guardrails for designers and software engineers, creating a foundation for developing effective and human-centered AI agent interactions. This study contributes to a structured foundation for future empirical studies on agentic AI in enterprise settings.
Kathrin Paimann, Elizângela Valarini, Sebastian Juhl· arXiv.org· 0 citations
Abstract Artificial Intelligence (AI) is typically considered an efficiency tool in user experience (UX). Far less attention is paid to its potential to structure reflection. The AI-supported UX Strategy Evaluator, created by German UPA’s UX Strategy Working Group, addresses this gap by posing questions to UX strategists and UX managers, rather than attempting to provide answers or best practices. This paper describes its design principles and development, grounded in the iterative development of a UX strategy checklist to identify which skills, information, and attitudes are relevant for strategic UX work. In addition, it describes a first exploration in a workshop with 13 participants at the conference Mensch & Computer 2025. Initial findings indicate that the question-based approach encourages participants to make assumptions explicit and fosters dialogue. The article concludes with areas for improvement and outlines future work.
Franziska Gronwald, Björn Rohles· i-com· 0 citations
Autonomous artificial intelligence (AI) agents can now log into a learning management system, read course materials, and complete unproctored, asynchronous assessed work end-to-end with no student involvement. We document that capability and trace its consequences for assessment validity. Three demonstrations on a live undergraduate course supply the evidence: two quiz completions, one in approximately 12 minutes, one in under 5, and a third in which the agent fabricated credible personal reflection for a discussion board. The wider public record includes at least 15 documented agent runs across three platforms and seven tools. We apply Kane’s argument-based validity framework: agent completion removes the attribution on which every inference in Kane’s chain depends. Everything downstream, from course grades to the evidence chains behind program review and accreditation, rests on support that is no longer there. The failure concerns validity rather than integrity: an institution can punish misconduct and still lack grounds for the scores it reports. Collective accreditor guidance addresses institutional uses of AI in evaluation and does not yet reach the agentic case. Audience data from the underlying conference session show attendees already recognizing both the vulnerability and the gap in institutional guidance. Polled attendees most often named online quizzes as agent-completable, with discussion-based work close behind. Majorities in both listings were working without settled written guidance. The response defended here is design rather than detection: four principles for verified human presence, low-effort changes faculty can adopt now, and the assurance levers assessment professionals already operate.
Stavros P. Hadjisolomou, R. El-Haddad· Intersection: A Journal at t...· 0 citations
Related blog posts
Microsoft Research Blog· microsoft.comJul 30, 2026
Computer-use AI agents struggle with multi-step workflows like email and customer support. Echoverse trains agents in realistic environments rather than simply providing more training tasks, helping them improve as the tasks, tests, and environments evolve. The post Echoverse: Deep, evolving environments for computer-use agents appeared first on Microsoft Research.
MIT News · Artificial Intelligence· news.mit.eduJul 14, 2026
Through research and entrepreneurship, Professor Devavrat Shah is helping to design methods that can handle constant decision-making using limited computational resources.
MIT News · Artificial Intelligence· news.mit.eduJun 3, 2026
The new ChartNet training dataset could improve the accuracy of vision-language models that help analyze business trends or interpret scientific figures.