Skip to content
Conference

An LLM-Augmented Framework for Embedded Firmware Development with Automated Validation and Hardware-in-the-Loop Testing

Aug 2026 · 2026 2nd International Conference on Electronic Information, Computer and Aerospace Remote Sensing (EICARS) · pp. 344-350 · 0 citations · 15 references

Abstract

The application of Large Language Models (LLMs) to embedded firmware development presents unique challenges: domain-specific register-level programming, stringent resource constraints, and the absence of structured validation mechanisms. Existing approaches using direct LLM-assisted coding suffer from hallucinated library functions, incorrect register configurations, and low compilation success rates, particularly for resource-constrained microcontrollers. This paper presents a three-layer framework that integrates DeepSeek-based code generation with an automated cross-compilation validation pipeline and hardware-in-the-loop (HIL) testing. The framework comprises an LLM integration layer with a local protocol proxy bridge, a three-stage code validation pipeline (static analysis, cross-compilation, and HIL testing), and a domain-specific prompt engineering module. Empirical evaluation across five benchmark tasks on ESP32-S3 and STM32F103 platforms demonstrates that the framework reduces development time by 53.7% on average compared to manual development (p < 0.001, Cohen's $\mathrm{d} = 2.34)$, while achieving 89.2% functional correctness—a 16.7 percentage point improvement over direct LLM-assisted coding $(\mathrm{p}<0.01$, Cohen's $\mathrm{d}=1.18)$. Error analysis reveals that hallucinated API calls are reduced by 77.2% and logic/timing errors by 74.1%. Compared to the recently proposed AutoEmbed system, our framework achieves comparable end-toend success rates (89.2% vs. 86.5%) while providing explicit cross-compilation and HIL validation stages absent from AutoEmbed's auto-programming approach. The framework leverages DeepSeek-V4-Flash for code generation, achieving substantial cost reduction compared to GPT-4-based alternatives based on published API pricing, making it economically viable for iterative embedded development workflows.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.