Sep 2026· ACM Transactions on Software Engineering and Methodology· 0 citations· 44 references
Advanced Malware Detection TechniquesSoftware Engineering Research
TL;DR
Hunk-Constrained Direct Preference Optimization is introduced, a training framework that unifies security hardening and functional correction in large language models and demonstrates that HPO achieves substantial security improvements—up to 28 percentage points—while preserving or enhancing functional correctness.
Abstract
Large Language Models (LLMs) have been widely applied in code generation tasks like code completion and automated development, demonstrating significant potential for improving coding efficiency. However, research has shown that LLM-generated code frequently contains security vulnerabilities, raising concerns about its reliability in production environments. To address these security issues, various mitigation approaches have been proposed, but these methods typically impact the LLM’s ability to generate functionally correct code, which may limit their practical application in real-world development environments. In this work, we address this problem through a key observation: security patches and functional bug fixes in real-world software exhibit structural similarities as small, localized modifications. This shared characteristic suggests that a unified learning model could address both objectives jointly. Building on this insight, we introduce HPO (Hunk-Constrained Direct Preference Optimization), a training framework that unifies security hardening and functional correction. Our framework features two key technical components: a novel segment-weighted preference optimization objective to focus learning on repair logic, and an automated data synthesis pipeline to provide high-quality training data. Experiments across multiple models and programming languages demonstrate that HPO achieves substantial security improvements—up to 28 percentage points—while preserving or enhancing functional correctness.
The first systematic study of model editing as a model-level hardening mechanism for secure code generation is conducted, evaluating 3 state-of-the-art editing methods across diverse LLM families and comparing them with CoSec, a representative inference-time approach, focusing on security, robustness, generalization, a...
Wei-Feng Sun, Quan-Jun Zhang, Yuchen Chen et al.· 0 citations
The findings suggest that hidden states are a robust and informative resource for estimating functional code correctness, supporting a two-stage workflow that combines response-level risk screening with targeted line-level prioritization.
Multi-SALLM, a benchmarking framework designed to systematically evaluate Large Language Models’ ability to generate secure code, reveals three key findings: functional correctness and security are closely related but not equivalent, and sampling strategy is a critical risk factor.
Mohammed Latif Siddiq, Noshin Ulfat, Nishat Raihan et al.· International Conference on...· 0 citations
Due to their black-box nature, LLMs suffer from limited explain- ability and a lack of determinism. Their usage cost can also rise, particularly with repetitive tasks on large codebases. To mitigate this, we conduct a novel empirical study targeting three domain- specific languages for transformation rules, namely Comb...
Axel Allain, Aymeric Blot, D. Khelladi et al.· 1 citation
This work evaluates Jev, a decision model that instead selects directly from a declared set of candidates and returns a probability for each, against six open-weight autoregressive models and a frontier proprietary model, and builds JevVibe, a diagnosis-guided repair agent that uses predicted CWE labels to repair code...
Arshak Rezvani, Sasha Behrouzi, Ahmad Sadeghi· 0 citations
The results indicate that functional correctness in code generation can be meaningfully improved without modifying the backbone architecture, by jointly optimizing how tasks are prompted, how the model is adapted, and how final outputs are selected.
Able to defeat top-ranked human players and more efficient than other models, the new system could help decision-makers in military maneuvers or business negotiations.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.