A Mutation-Driven Trustworthiness Evaluation Method for LLM-Based Airborne Code Generation
Large language models (LLMs) have demonstrated significant potential for airborne embedded code generation, yet existing evaluation methods based primarily on the Pass@k metric focus on functional correctness and fail to adequately assess robustness degradation under semantic perturbations or identify domain-specific c...