2026· IEEE Transactions on Reliability· Vol 75, pp. 3178-3192· 0 citations· 52 references
Abstract
Large language models (LLMs) demonstrate strong capabilities in code-related tasks, however their effectiveness in software vulnerability detection (SVD) remains poorly understood due to inadequate evaluation frameworks. Existing benchmarks suffer from training data contamination, isolated function evaluation without cross-component context, lack of strict vulnerable-patched pairing, and binary classification without hierarchy-aware common weakness enumeration (CWE) assessment, which prevents a reliable measurement of security reasoning versus pattern matching on leaked data. We introduce LLMKernelBench, a rigorous benchmark comprising 417 real-world Linux kernel vulnerabilities across 74 CWE types, split into a primary benchmark dataset (PBD; 314 samples, $\leq$2024) and a leakage free dataset (LFD; 103 samples, 2025 post-cutoff), with context-aware evaluation at three granularity levels and hierarchy-aware metrics quantifying semantic proximity in misclassifications. We evaluate seven LLM spanning code-specialized and general-purpose architectures. Binary vulnerability detection is near-random ($\sim\!\! 50\%$ accuracy) and strongly biased: some models label $>75\%$ of samples as vulnerable, while others mostly label them as nonvulnerable. CWE prediction is effectively unusable, with an average Top-1 accuracy of 1.4% (best: 3.3%) and a 12.3% hierarchy proximity score, providing little reliable exact or taxonomy-level signal. On the leakage-free 2025 split, binary accuracy remains near-random and robustness to multifile abstraction is model-specific rather than tied to specialization, with code-specialized and general-purpose models degrading by 2.6% and 4.5% on average, respectively. The micro-to-macro CWE-accuracy gap is larger on LFD (6.6 points) than on PBD (0.9 points), which is consistent with sensitivity to class frequency but does not identify an internal model mechanism.
JavaScript powers approximately 98.8% of all websites, making vulnerabilities in its code a significant security risk, yet existing detection approaches such as Static Application Security Testing (SAST) tools often fail to identify many real-world vulnerabilities when applied to isolated code snippets. This paper pres...
Manit Kaushik, Ishir Bhardwaj, Pranav Gupta et al.· 0 citations
This work presents SCRIPTIOC-BENCH, a benchmark for measuring static IOC extraction capability on real-world malicious scripts, and evaluates a broad range of proprietary and open-weight LLMs, showing that IOC recovery without execution remains challenging across model scales.
Hanna Kim, Jian Cui, Minkyoo Song et al.· 0 citations
VICBench enables robust evaluation of vulnerability detection approaches and shows that state-of-the-art algorithms V-SZZ and LLM4SZZ achieve only 33.3%-40.1% F1, confirming that using existing approaches still entails significant manual effort.
Jin Lu, Xuening Han, Yan Zhong et al.· 0 citations
Despite the dominance of Transformer-based models in software vulnerability detection, the extent to which their learned security logic generalizes across different programming languages remains a critical open question. To address this, we propose a comprehensive evaluation framework organized into three phases spanni...
Nhien Huu Dinh, Chau The, Thai Hung Van et al.· International Conference on...· 0 citations
This thesis rebuilds Real-Vul through a controlled Code Property Graph pipeline, measure and remove a 37% content leak intrinsic to whole-codebase sampling, and train a relational graph neural network with a disciplined class-imbalance recipe: focal loss, class-aware undersampling, and fine-tuning of a GraphCodeBERT no...
Large language models (LLMs) embedded in enterprise workflows cannot structurally distinguish legitimate instructions from adversarial ones in the same token stream, making prompt injection OWASP's top LLM risk for two consecutive editions a persistent threat across direct and indirect vectors. This paper presents Prom...
Fatimah Alhamzawi· Al-Noor Journal of Engineeri...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.