Skip to content
Review Open access

Automated Testing Techniques for Enterprise Software Systems with GenAI Integration

Sep 2026 · International Journal of Advanced Artificial Intelligence Research · Vol 03, pp. 74-102 · 0 citations

TL;DR

An automated testing framework that combines metamorphic testing and property-based testing to validate enterprise GenAI applications at scale is developed, providing a practical path to deployment readiness, regulatory traceability, and sustained operational trust.

Abstract

The rapid adoption of Generative Artificial Intelligence (GenAI) in enterprise workflows, from customer support to document processing, procurement and knowledge retrieval, has revealed a fundamental gap in conventional quality assurance practice. GenAI components are probabilistic, context-sensitive, dependent on retrieval corpora, tool integrations and evolving model versions, unlike deterministic software. Automated testing frameworks based on the assumption of stable input-output mappings are no longer sufficient to ensure the reliability, safety and compliance of these systems. In regulated and high-stakes enterprise settings, the lack of a scalable, CI/CD-friendly testing approach creates an unacceptable risk. In this paper, we develop an automated testing framework that combines metamorphic testing and property-based testing to validate enterprise GenAI applications at scale. Metamorphic testing identifies inconsistencies and emergent faults between related input transformations without needing to define the expected outputs a priori, directly addressing the oracle problem faced by generative systems. Property-based testing is an alternative approach that creates a range of test cases from business rules, domain invariants and constraint specifications. Simultaneously, the framework evaluates hallucination rates, accuracy of retrieval grounding, correctness of tool calls, policy compliance, and prompt regression for RAG pipelines and agentic workflows. To demonstrate how the proposed framework's metrics would be applied and evaluated in practice, this paper additionally presents an illustrative case scenario across four representative enterprise workflows customer support triage, procurement approvals, code review assistance, and knowledge-base Q&A conceptually modeled on published, evidence-driven quality-gate approaches for LLM applications. In this constructed 24-week quasi-experimental scenario, automated testing gates are illustrated as producing a 6.7 percentage point increase in task success rate (Cohen's d = 1.52), a 27.5% decrease in escalation rate, and a 31.0% decrease in rework rate. Business error cost per 1,000 workflow instances is illustrated as decreasing by 33.4%, and incorrect tool actions by 38.5%. User satisfaction is illustrated as increasing by 0.34 points on a 5-point scale. These illustrative results are shown to hold up under difference-in-differences and segmented-regression analyses, demonstrating how a framework's effectiveness could be assessed beyond simple pre/post comparison rather than reporting outcomes of a completed deployment. Taken together, the proposed framework and illustrative demonstration position metamorphic and property-based testing as scalable, CI/CD-compatible quality assurance techniques for enterprise GenAI systems, providing a practical path to deployment readiness, regulatory traceability, and sustained operational trust.

Read PDF

Similar papers

Review Open access Sep 2026

Artificial Intelligence-Driven Software Test Automation: A Comprehensive Survey

Software testing accounts for a significant proportion of resource consumption within the software development lifecycle. Traditional automated testing methods rely on predefined scripts and deterministic logic, and have inherent limitations when addressing the dynamic complexities of modern software systems. Artificia...

Jia-Lei Chen · 0 citations
Open access Sep 2026

Automated Quality Assurance in Web Applications Using Model-Based Testing and Cypress with AI-Assisted Test Generation

Modern web applications require robust, scalable and efficient QA methodologies due to their increasing complexity. Dynamic user interfaces, frequent deployments, and changing functional requirements many times can not be well addressed by manual testing or scripting tests approaches. In this paper, we outline an autom...

Bhuvan Chandra Kasarapu · 0 citations
Preprint Aug 2026

The Specification Paradox: Rethinking Requirements Engineering in the Age of AI

The growing adoption of Large Language Models (LLMs) in Software Engineering has reinforced the expectation that coding activities can be largely automated. However, this perception may represent yet another historical search for a solution capable of eliminating the inherent challenges of software development. This ar...

T. Sirqueira, Jessica Faciroli · 1 citation
Open access Aug 2026

DEVELOPING AN AUTOMATED FUNCTIONAL TESTING FRAMEWORK - A PRACTICAL CASE STUDY

Test automation addresses limitations such as time consuming, error prone, and difficult to reproduce at a scale by enabling systematic, efficient, and repeatable validation of software functionality across multiple deployments. This project presents an analysis of the client’s CRM system, with the primary objective of...

Rajat Sharma, Shahid Ali · 0 citations
#software testing Conference Open access Sep 2026

AI-Assisted Software Testing: Opportunities and Challenges

It is argued that AI is unlikely to fully replace human testers in the near future and should be used as an assistant that supports human judgment in software quality assurance.

Ming-Yang Peng · 0 citations
#software testing Preprint Sep 2026

A Large-Scale Empirical Study of Quality Assurance Practices and Gaps in AI Agents

A large-scale empirical study of quality assurance (QA) practices in 157 open-source LLM-based agent projects with at least 100 GitHub stars highlights the need to move beyond feature-level testing toward systematic end-to-end validation that ensures agent workflows remain within intended boundaries when interacting wi...

Wu-Yang Dai, Moses Openja, Jiho Shin et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.