Skip to content

JevVibe: Efficient Classification-Guided Secure Code Generation

Sep 2026 · 0 citations · 13 references
Computer Science

TL;DR

This work evaluates Jev, a decision model that instead selects directly from a declared set of candidates and returns a probability for each, against six open-weight autoregressive models and a frontier proprietary model, and builds JevVibe, a diagnosis-guided repair agent that uses predicted CWE labels to repair code generated by Qwen2.5-Coder-32B-Instruct.

Abstract

Large language models can generate functionally correct code that still contains security weaknesses, motivating repair pipelines that first diagnose a weakness type before deciding how to fix it. The Common Weakness Enumeration (CWE) provides a standardized vocabulary for such diagnoses, but asking an autoregressive language model to generate a CWE label and extracting it from the response raises questions about output validity, speed, and cost, as well as accuracy. We evaluate Jev, a decision model that instead selects directly from a declared set of candidates and returns a probability for each, against six open-weight autoregressive models and a frontier proprietary model, GPT-5.6-Sol, on a controlled 50-way CWE classification task over 1,916 CyberSecEval benchmark examples. Jev outperforms all six open-weight baselines on every classification and ranking metric, while its comparison with GPT-5.6-Sol depends on the metric: GPT-5.6-Sol achieves higher Top-1 accuracy and Macro-F1, whereas Jev achieves higher Top-3 and Top-5 accuracy and a nearly identical MRR, at $6.27\times$ lower median API latency and $55.9\times$ lower estimated API cost. We further build JevVibe, a diagnosis-guided repair agent that uses predicted CWE labels to repair code generated by Qwen2.5-Coder-32B-Instruct. With Jev providing the diagnosis, the agent increases the detector-measured security pass rate from 63.5% before repair to 70.7%, compared with 66.1% for LLM-guided repair. These results show that JevVibe is effective at improving the security of generated code, with Jev providing reliable and efficient CWE classification.

View source

Similar papers

Hunk-Constrained DPO: Segment-Level Optimization for Secure and Correct LLM Code Generation

Hunk-Constrained Direct Preference Optimization is introduced, a training framework that unifies security hardening and functional correction in large language models and demonstrates that HPO achieves substantial security improvements—up to 28 percentage points—while preserving or enhancing functional correctness.

Qian-Shuo Huang, Xin Yin, Xin-Rui Li et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Probing the Prefill: Detecting Code Vulnerabilities via Latent Activations

LLM-based code generation is now embedded in mission-critical pipelines, but defenses against vulnerable output remain post-hoc -- static analyzers, fine-tuned classifiers, or an LLM judge that screen completed code, ignoring the generating model's own internal state. We test a narrower, directly measurable question: w...

Alizishaan Khatri · 0 citations
Preprint Aug 2026

Detecting Contaminated Code-Generation Prompt Batches via Influence Functions

Large language models (LLMs) are increasingly used for code generation, yet they remain vulnerable to prompts that elicit insecure implementations. Existing defenses typically rely on predefined threat models or known vulnerability patterns, limiting their effectiveness against novel attacks. We propose CodeSIFT, a thr...

Francesco Quinzan, Noor Munir, Yi-Shun Lu et al. · 0 citations
#computer vision Preprint Aug 2026

CodeAssay: A Multi-Metric Benchmark with Audited Ground Truth for LLM Code Generation

These findings show that reliable evaluation of LLM-generated code requires validated ground truth, protected tests, and multiple explicitly interpreted measures, and that CodeAssay provides a reproducible basis for evidence-based model evaluation in AI-augmented software development.

Shahbaz Siddeeq, Muhammad Waseem, Umar Subhan Malhi et al. · 0 citations
Preprint Open access Aug 2026

Understanding and Improving Model Editing for Secure Code Generation

The first systematic study of model editing as a model-level hardening mechanism for secure code generation is conducted, evaluating 3 state-of-the-art editing methods across diverse LLM families and comparing them with CoSec, a representative inference-time approach, focusing on security, robustness, generalization, a...

Wei-Feng Sun, Quan-Jun Zhang, Yuchen Chen et al. · 0 citations
Open access Aug 2026

Multi-SALLM: a multilingual security assessment of generated code

Multi-SALLM, a benchmarking framework designed to systematically evaluate Large Language Models’ ability to generate secure code, reveals three key findings: functional correctness and security are closely related but not equivalent, and sampling strategy is a critical risk factor.

Mohammed Latif Siddiq, Noshin Ulfat, Nishat Raihan et al. · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.