A Mutation-Driven Trustworthiness Evaluation Method for LLM-Based Airborne Code Generation
Abstract
Large language models (LLMs) have demonstrated significant potential for airborne embedded code generation, yet existing evaluation methods based primarily on the Pass@k metric focus on functional correctness and fail to adequately assess robustness degradation under semantic perturbations or identify domain-specific compliance deficiencies in aviation applications. To address the trustworthiness evaluation challenge for airborne code generation, this paper proposes a mutation-driven trustworthiness evaluation method called MuTEval. This method integrates the TREAT evaluation framework, mutation testing, and expert review. MuTEval establishes a three-layer trustworthiness indicator system comprising: (1) the functional correctness dimension, including pass@k, test pass rate (TPR); (2) a robustness dimension, including consistency pass rate (CPR), and performance degradation rate (PDR), and mutation sensitivity (MS); and (3) a domain applicability dimension, encompassing safety compliance, coding standard conformance, real-time guarantee, maintainability. Based on the AvionicsEval, a benchmark consisting of 50 airborne embedded C code generation tasks, four models-DeepSeek-Coder-V2, Code LLaMA-34B, Qwen2.5-Coder-32B, and GPT-4o-are evaluated. Experimental results show that 18%-28% of tasks that pass conventional tests fail under semantic mutations, while domain applicability scores decrease by up to 24.6 percentage points relative to functional correctness scores. These findings demonstrate that MuTEval surpasses traditional Pass@k-based evaluation by providing a more comprehensive and effective assessment, enabling identification of domain-specific deficiencies in LLM-generated airborne code.