From Voice to Robot: Evaluating Large Language Models for Industrial Programming with HARPA
Abstract
Recent advances in generative artificial intelligence and large language models (LLMs) have increased the feasibility of translating natural-language instructions into executable robot programs, reducing the technical barrier separating shop-floor operators from industrial robotic programming. However, current evaluation practices remain insufficient, as they frequently privilege textual similarity, syntactic validity, or isolated execution success without assessing whether generated programs are reliable, auditable, accountable, precise, and aligned with human intent. This limitation is particularly relevant in Industry 5.0, where robotic systems must support transparent, human-centred, and trustworthy collaboration. This study presents HARPA, a multidimensional evaluation framework for LLM-generated robot code, structured around five dimensions: Human Alignment, Accountable, Reliable, Precise, and Auditable. The framework was validated through a natural-language-to-code-to-simulation pipeline in which Portuguese instructions were converted into three industrial robot programming languages: KRL, RAPID, and URScript. Three models were evaluated under two acoustic conditions: optimal input and simulated factory noise. Across 900 automated executions, the pipeline achieved an overall success rate of 96.1% (95% Wilson score confidence interval: 94.6%–97.2%), with 96.4% under optimal conditions and 95.8% under noise. The difference between acoustic conditions was not statistically significant in the aggregated success analysis ( p = 0.605), whereas model-level results showed higher Reliable scores for proprietary models than for the open-source baseline. The results show that LLMs can support controlled robot programming tasks while also identifying directions for further development, particularly regarding Precise scores, scenario complexity, and adaptation to broader industrial use. HARPA contributes an execution-oriented, human-centred evaluation layer for safer AI-assisted robotic programming.