While LLM-based multi-agent systems show potential for large-scale software development, successful integration requires addressing challenges such as memory limitations, hallucinations, and code smells, alongside a practitioner-centric perspective.
Z. Rasheed, Muhammad Waseem, Kai-Kristian Kemell et al.· 31 citations
The proposed LLM-based multi-agent system automates qualitative data analysis process, creating opportunities for researchers and practitioners, and future improvements focus on enhancing multilingual performance and integrating continuous expert feedback.
Z. Rasheed, Muhammad Waseem, Aakash Ahmad et al.· arXiv.org· 41 citations
This work designs and operationalizes a governance-aware, multi-tenant AI sandbox that supports structured experimentation and produces reusable evaluation evidence across stakeholders and yields lessons learned and practical considerations that inform deployment and future evolution of governance-aware sandbox platfor...
Muhammad Waseem, M. Islam, Md Nasir Uddin Shuvo et al.· arXiv.org· 0 citations
This work presents the design and implementation of a governance-aware, multi-tenant AI sandbox for structured experimentation and the generation of reusable evaluation evidence across projects and stakeholder groups.
Muhammad Waseem, M. Islam, Md Nasir Uddin Shuvo et al.· 0 citations
A Multi-Vocal Literature Review is conducted, combining insights from both academia and industry, including peer-reviewed studies and grey literature to systematically synthesize and analyze existing knowledge on LLM-based multi-agent systems for code generation.
Z. Rasheed, Muhammad Waseem, Kai-Kristian Kemell et al.· arXiv.org· 2 citations
This study provides the first large-scale empirical comparison of agentic frameworks for reasoning-intensive software engineering tasks and shows that framework selection should prioritize orchestration quality, especially memory control, failure handling, and cost management.
Z. Rasheed, Malik Abdul Sami, Muhammad Waseem et al.· arXiv.org· 1 citation
The vision is to leverage the capabilities of multiple GPT agents to contribute to SE tasks and to propose an initial road map for future work, arguing that multiple G PT agents can perform creative and demanding tasks far beyond coding and debugging.
Z. Rasheed, Muhammad Waseem, Kai-Kristian Kemell et al.· XP Workshops· 34 citations· ⚡2
This research highlights the heightened threats to data integrity and stakeholder trust in these evolving ecosystems through an intensive examination of the literature, initiating a pioneering discourse emphasizing fostering a foundation for developing secure and trustworthy Liquid AI environments.
M. Agbese, Niko Mäkitalo, Muhammad Waseem et al.· IoT· 6 citations· ⚡1
Carbon-Aware Governance Gates (CAGG), an architectural extension that embeds carbon budgets, energy provenance, and sustainability-aware validation orchestration into human-AI governance layers, is proposed.
M. Abbasi, T. Mikkonen, Petri Ihantola et al.· 2026 IEEE 23rd International...· 0 citations
Abstract. Carbon dioxide (CO2) emissions from industrial activities remain one of the greatest contributors to global climate change. Hollow fiber membranes (HFMs) have emerged as a promising technology for post-combustion CO2 separation owing to their high surface-area-to-volume ratio and scalability. This work focuse...
Muhammad Waseem· Materials Research Proceedin...· 0 citations
End-to-end weather forecasting systems produce skillful global gridded and station forecasts directly from raw Earth observations, replacing the numerical weather prediction pipeline, including data assimilation, at a fraction of its cost. These systems are deterministic and issue no uncertainty. Here we render the Aar...
Rodrigo Almeida, Noelia Otero, Jost Arndt et al.· 0 citations
Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity. The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.