TAILOR broadens the coverage of automated CVE reproduction and provides auditable evidence for vulnerability diagnosis and defense and ablation experiments show that the two control levels respectively mitigate execution-path mismatch and missing Web prerequisite state.
Abstract
Growing vulnerability disclosure and widespread software reuse increase security teams'need for reproducible evidence to diagnose vulnerabilities, validate patches, and build regression tests. Producing such evidence at scale requires automated end-to-end CVE reproduction. Existing methods typically process different CVEs through a uniform pipeline, but differences in runtime form, trigger interfaces, and prerequisite state impose different execution requirements on individual stages, making fixed workflows difficult to adapt to diverse reproduction needs. To address this problem, we present TAILOR, a type- and state-aware multi-agent framework specialized for complex vulnerability reproduction. TAILOR converts static vulnerability information into auditable reproduction evidence and packages reconstructed environments and trigger evidence into reproduction artifacts. Its first-level type-aware mechanism adaptively matches each vulnerability to an execution path. Within the Web path, its second-level state-aware mechanism constructs the required prerequisite state before exploitation, decouples prerequisite-state construction from core vulnerability triggering, and shares execution constraints across exploitation and verification. We construct a dataset of 200 CVEs with an emphasis on cases with complex execution requirements. TAILOR successfully reproduces 59.24\% of Web vulnerabilities and 44.19\% of traditional vulnerabilities. Further ablation experiments show that the two control levels respectively mitigate execution-path mismatch and missing Web prerequisite state. Overall, TAILOR broadens the coverage of automated CVE reproduction and provides auditable evidence for vulnerability diagnosis and defense.
Continuous integration and continuous deployment pipelines produce security-relevant records across source, build, test, package, deployment, and runtime environments. However, these records remain fragmented across logs, scanner findings, policy decisions, identity events, artefact metadata, and runtime alerts, limiti...
Sabbir M. Saleh, N. Madhavji, John Steinbacher· Proceedings of the ACM/IEEE...· 0 citations
BUGSTONE-E2E, a framework that transforms vulnerability history into executable detection rules and validates their findings, demonstrates that CVE history can be turned into an executable workflow, transforming past vulnerabilities into reproducible detection and repair.
Qiu-Shi Wu, Kevin Eykholt, Youngja Park et al.· 0 citations
VICBench enables robust evaluation of vulnerability detection approaches and shows that state-of-the-art algorithms V-SZZ and LLM4SZZ achieve only 33.3%-40.1% F1, confirming that using existing approaches still entails significant manual effort.
Jin Lu, Xuening Han, Yan Zhong et al.· 0 citations
Reliable learning-based vulnerability detection requires high-quality labels, yet datasets built from vulnerability-fixing commits may label functions as vulnerable simply because they were changed by a security patch. We present VulValidate, a framework that uses LLM agents to coordinate dynamic analysis tools and con...
Large language model (LLM) agents are increasingly evaluated on cybersecurity tasks such as vulnerability reproduction, exploitation, and patching. However, existing cybersecurity benchmarks predominantly operate under a post-environment evaluation paradigm, i.e., handing the agent source code, a container, or an execu...
Liang He, Sheng Wu, Hao-Miao Hao et al.· 0 citations