Evaluation of large language models in api testing using open API documentation
Abstract
Conventional automated REST API testing approaches often depend on rule-based logic, extensive configuration, or source-code access, which limits their adaptability in rapidly evolving development environments. Recent advances in Large Language Models (LLMs) offer new possibilities for automating API test generation through their ability to interpret structured documentation such as OpenAPI specifications. However, systematic empirical evaluation of LLM performance in API testing remains limited. This study evaluates the effectiveness of multiple state-of-the-art LLMs in generating executable REST API test cases from OpenAPI documentation. Four LLMs are evaluated using four prompting strategies: zero-shot, one-shot, few-shot, and chain-of-thought prompting. Thirteen public REST APIs from diverse domains and varying levels of complexity are selected to construct a representative evaluation dataset. Model performance is measured using a multi-dimensional evaluation framework, including coverage metrics, an oracle quality metric, and a test validity metric. Descriptive analysis is complemented by non-parametric statistical methods, including the Kruskal-Wallis H test, Dunn’s post-hoc test, Spearman’s rank correlation, and the Wilcoxon signed-rank test. The results indicate that LLMs are capable of generating valid and executable API test suites directly from OpenAPI specifications, with statistically significant performance differences across models and prompting strategies. Based on these findings, the study also proposes a generalized prompt template that balances reasoning quality and practical usability for exploratory API testing. The findings demonstrate the practical feasibility of LLM-driven API testing and provide empirical guidance for model selection and prompt design for exploratory testing.