Back to feed
Open access

Can Language Models Generate Secure Terraform Code? A Security-Focused Benchmark Using Static Analysis

Jul 2026 · Anais do I Simpósio de Infraestrutura Digital/Nuvem para Pesquisa (Pesquisa@Nuvem 2026) · 0 citations · 15 references

Abstract

The widespread adoption of Infrastructure-as-Code (IaC) has made cloud misconfiguration a critical security concern, while Large Language Models (LLMs) and Small Language Models (SLMs) have been redefining the programming process. We present an empirical benchmark evaluating whether LLMs and SLMs can generate security-compliant AWS Terraform configurations. Our automated pipeline integrates Checkov and Trivy into a GitLab CI/CD workflow across four Amazon S3 scenarios, evaluating six models under three prompt strategies of increasing security specificity. Security compliance improves consistently with prompt detail, though no model achieves full compliance in any configuration. Our findings suggest that prompt design is a critical factor, highlighting the need for a proper pipeline for developing and validating LLM-assisted secure IaC generation. All artifacts are publicly available.

Read PDF