Software systems increasingly embed a large language model in features that must satisfy a numeric output constraint, that is, a requirement expressible as a number or an interval and checkable by code, such as a target word count or a target readability grade band. Because such a model is non-deterministic, is configu...
Quan Zhou, Shahbaz Siddeeq, Mika Saari et al.· 0 citations
A structured community discussion at the First International Workshop on Empirical Prompt Engineering for Software Engineering (PROMPT-SE) discussed current prompting practices, challenges to their adoption and evaluation, and future directions for integrating prompt engineering into software development.
Vincenzo De Martino, Giovanna Broccia, Fabiano Pecorelli et al.· 0 citations
A large language models based multi-agent system enables precise task execution and inter-agent collaboration, addressing the challenges of refactoring in functional programming.
Shahbaz Siddeeq, Z. Rasheed, Malik Abdul Sami et al.· arXiv.org· 1 citation
An LLM-based multi-agent system that autonomously upgrades legacy web applications to the latest versions and maintains context across tasks and agents, improving solution quality over the base model in some cases is proposed.
Valtteri Ala-Salmi, Z. Rasheed, Malik Abdul Sami et al.· International Conference on...· 4 citations
This is one of the first reviews to integrate peer-reviewed and grey literature on vibe coding under a single documented protocol and is strongest for prototyping and user-interface work and weakest for production, data-intensive, and safety-critical use, and tool visibility does not imply effectiveness.
Shahbaz Siddeeq, Muhammad Waseem, Kai-Kristian Kemell et al.· arXiv.org· 0 citations
Overall, the results suggest that epic-organized generation can improve perceived Gherkin quality while maintaining comparable semantic coverage, although broader replication is needed before generalizing this finding.
Shahbaz Siddeeq, M. Abbasi, Jussi Rasku et al.· arXiv.org· 0 citations
Results highlight the ability of LLM-based multi-agent in managing refactoring tasks targeted toward functional programming paradigms and hint that LLM-based multi-agent systems integration into the refactoring of functional programming languages can enhance maintainability and support automated development workflows.
Shahbaz Siddeeq, Muhammad Waseem, Z. Rasheed et al.· International Conference on...· 5 citations
These findings show that reliable evaluation of LLM-generated code requires validated ground truth, protected tests, and multiple explicitly interpreted measures, and that CodeAssay provides a reproducible basis for evidence-based model evaluation in AI-augmented software development.
Shahbaz Siddeeq, Muhammad Waseem, Umar Subhan Malhi et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.