Skip to content

LLMKernelBench: Benchmarking Large Language Models on Software Vulnerability Detection in Linux Kernel

2026 · IEEE Transactions on Reliability · Vol 75, pp. 3178-3192 · 0 citations · 52 references

Abstract

Large language models (LLMs) demonstrate strong capabilities in code-related tasks, however their effectiveness in software vulnerability detection (SVD) remains poorly understood due to inadequate evaluation frameworks. Existing benchmarks suffer from training data contamination, isolated function evaluation without cross-component context, lack of strict vulnerable-patched pairing, and binary classification without hierarchy-aware common weakness enumeration (CWE) assessment, which prevents a reliable measurement of security reasoning versus pattern matching on leaked data. We introduce LLMKernelBench, a rigorous benchmark comprising 417 real-world Linux kernel vulnerabilities across 74 CWE types, split into a primary benchmark dataset (PBD; 314 samples, $\leq$2024) and a leakage free dataset (LFD; 103 samples, 2025 post-cutoff), with context-aware evaluation at three granularity levels and hierarchy-aware metrics quantifying semantic proximity in misclassifications. We evaluate seven LLM spanning code-specialized and general-purpose architectures. Binary vulnerability detection is near-random ($\sim\!\! 50\%$ accuracy) and strongly biased: some models label $>75\%$ of samples as vulnerable, while others mostly label them as nonvulnerable. CWE prediction is effectively unusable, with an average Top-1 accuracy of 1.4% (best: 3.3%) and a 12.3% hierarchy proximity score, providing little reliable exact or taxonomy-level signal. On the leakage-free 2025 split, binary accuracy remains near-random and robustness to multifile abstraction is model-specific rather than tied to specialization, with code-specialized and general-purpose models degrading by 2.6% and 4.5% on average, respectively. The micro-to-macro CWE-accuracy gap is larger on LFD (6.6 points) than on PBD (0.9 points), which is consistent with sensitivity to class frequency but does not identify an internal model mechanism.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Exploring Automated Vulnerability Identification in JavaScript Code Using Large Language Models

JavaScript powers approximately 98.8% of all websites, making vulnerabilities in its code a significant security risk, yet existing detection approaches such as Static Application Security Testing (SAST) tools often fail to identify many real-world vulnerabilities when applied to isolated code snippets. This paper pres...

Manit Kaushik, Ishir Bhardwaj, Pranav Gupta et al. · 0 citations
#artificial intelligence Preprint Sep 2026

SCRIPTIOC-BENCH: A Benchmark for Recognizing Actionable Threat Intelligence from Script-Based Malware using LLMs

This work presents SCRIPTIOC-BENCH, a benchmark for measuring static IOC extraction capability on real-world malicious scripts, and evaluates a broad range of proprietary and open-weight LLMs, showing that IOC recovery without execution remains challenging across model scales.

Hanna Kim, Jian Cui, Minkyoo Song et al. · 0 citations
Preprint Aug 2026

VICBench: A Multi-Language Benchmark for Code Vulnerability Detection

VICBench enables robust evaluation of vulnerability detection approaches and shows that state-of-the-art algorithms V-SZZ and LLM4SZZ achieve only 33.3%-40.1% F1, confirming that using existing approaches still entails significant manual effort.

Jin Lu, Xuening Han, Yan Zhong et al. · 0 citations
Conference Aug 2026

An Empirical Study on the Transferability of Transformer-Based Models for Software Vulnerability Detection

Despite the dominance of Transformer-based models in software vulnerability detection, the extent to which their learned security logic generalizes across different programming languages remains a critical open question. To address this, we propose a comprehensive evaluation framework organized into three phases spanni...

Nhien Huu Dinh, Chau The, Thai Hung Van et al. · 0 citations
Open access

Recovering vulnerability detection on realistically-distributed code

This thesis rebuilds Real-Vul through a controlled Code Property Graph pipeline, measure and remove a 37% content leak intrinsic to whole-codebase sampling, and train a relational graph neural network with a disciplined class-imbalance recipe: focal loss, class-aware undersampling, and fine-tuning of a GraphCodeBERT no...

Brian Kade Betterton · 0 citations
Open access Aug 2026

Real-Time Detection and Mitigation of Prompt Injection Attacks in LLM-Integrated Enterprise Systems

Large language models (LLMs) embedded in enterprise workflows cannot structurally distinguish legitimate instructions from adversarial ones in the same token stream, making prompt injection OWASP's top LLM risk for two consecutive editions a persistent threat across direct and indirect vectors. This paper presents Prom...

Fatimah Alhamzawi · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.