Skip to content

Category

climate science

381 papers

#computer vision Apr 2024

Large Language Model Evaluation Via Multi AI Agents: Preliminary results

A novel multi-agent AI model is introduced that aims to assess and compare the performance of various LLMs, and initial results indicate that the GPT-3.5 Turbo model's performance is comparatively better than the other models.

Z. Rasheed, Muhammad Waseem, Kari Systä et al. · 23 citations
#computer vision Review Jan 2025

Large Language Models for Code Generation: The Practitioners Perspective

This work proposes and develops a multi-model unified platform to generate and execute code based on natural language prompts and presents practitioners feedback and insights into the use of LLMs in software development, including their strengths and weaknesses, key aspects overlooked by benchmarks and metrics.

Z. Rasheed, Muhammad Waseem, Kai-Kristian Kemell et al. · 18 citations · ⚡2
#computer vision Review Feb 2025

LLM-Generated Microservice Implementations from RESTful API Definitions

A system that uses Large Language Models (LLMs) to automate the API-first development of RESTful microservices and assists in creating OpenAPI specification, generating server code from it, and refining the code through a feedback loop that analyzes execution logs and error messages is presented.

Saurabh Chauhan, Z. Rasheed, Malik Abdul Sami et al. · 16 citations · ⚡1

VAPU: System for Autonomous Legacy Code Modernization

An LLM-based multi-agent system is indicated that an LLM-based multi-agent system is a capable solution to update components of a legacy application autonomously.

Valtteri Ala-Salmi, Z. Rasheed, Malik Abdul Sami et al. · 3 citations
#computer vision Jul 2025

Assessing Small Language Models for Code Generation: An Empirical Study with Benchmarks

This study presents a comprehensive empirical evaluation of 20 open-source Small Language Models and reveals that several compact SLMs achieve competitive results while maintaining a balance between performance and efficiency, making them viable for deployment in resource-constrained environments.

Mahade Hasan, Muhammad Waseem, Kai-Kristian Kemell et al. · 16 citations

Autonomous Legacy Web Application Upgrades Using a Multi-Agent System

An LLM-based multi-agent system that autonomously upgrades legacy web applications to the latest versions and maintains context across tasks and agents, improving solution quality over the base model in some cases is proposed.

Valtteri Ala-Salmi, Z. Rasheed, Malik Abdul Sami et al. · 4 citations
#computer vision Review Dec 2025

Vibe Coding in Practice: Flow, Technical Debt, and Guidelines for Sustainable Use

This article analyzes the flow-debt tradeoffs associated with VC and identifies and explains how current model, platform, and hardware limitations contribute to these issues, and proposes countermeasures to address them, informing research and practice towards more sustainable VC approaches.

Muhammad Waseem, Aakash Ahmad, Kai-Kristian Kemell et al. · 4 citations

From tech blogs

See all →
Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.