The Role of Expert Opinions and Peer Review in Engineering Next-Generation AI-Enabled Systems
Artificial intelligence (AI) has become a key enabling technology for business organisations, increasingly embedded in production systems and decision-making workflows. AI models and tools in these settings must satisfy diverse requirements, such as accuracy, traceability, explainability, fairness, and trust, while meeting compliance obligations, service-level agreements, and governance frameworks. However, the non-deterministic and evolving nature of modern AI models challenges traditional validation approaches based primarily on quantitative metrics. This paper investigates how techniques from empirical software engineering, particularly expert opinions and peer review, can support AI-enabled systems engineering. The paper presents an experience report grounded in a literature review and an experiment. The experiment is based on a precisely specified, complex problem that analyses the responses of contemporary AI models when prompted with the given problem. It examines how expert judgment and peer review can improve the plausibility and correctness of model outputs. This study proposes a lifecycle model that treats problem and ground-truth specifications, expert judgments, and review rationales as first-class artefacts. The paper concludes by discussing challenges, limitations, and opportunities for incorporating these techniques into AI engineering, particularly for next-generation models and tools.