Toward Dynamic and Risk-Aware Evaluation of Cybersecurity LLMs: A Survey and the RIRAG Framework
Abstract
The rapid adoption of large language models (LLMs) in cybersecurity has created a growing need for evaluation methods that reflect operational risk rather than isolated language capability. Existing cybersecurity benchmarks assess useful dimensions such as factual knowledge, vulnerability analysis, secure coding, penetration testing, and threat intelligence reasoning, but many remain limited by static datasets, weak diagnostic granularity, limited adversarial testing, and insufficient attention to human-AI decision-making. This survey analyzes recent LLM cybersecurity benchmarks through three evaluation paradigms: knowledge-oriented, task-oriented, and holistic evaluation. From this analysis, we identify five recurring gaps between benchmark performance and real-world cybersecurity risk: knowledge-reasoning mismatch, capability-risk separation, limited failure attribution, staticity and contamination, and adversarial fragility. To address these gaps, we introduce the Reflective and Iterative Retrieval-Augmented Generation (RIRAG) framework, a dynamic and risk-aware evaluation architecture for cybersecurity LLMs. RIRAG combines continuously updated cybersecurity knowledge, retrieval- and generation-specific metrics, diagnostic logging, independent evaluation, adversarial red teaming, operational risk scoring, and human-AI teaming assessment. The framework is specified through design principles, formal components, implementation guidance, and illustrative case studies. Rather than presenting RIRAG as a validated production system, this article defines a falsifiable empirical validation protocol for future study. The central contribution is a survey-grounded reference architecture for evaluating cybersecurity LLMs as evolving, adversarially exposed, and human-interactive systems.