Skip to content
Preprint

Post-Hoc Attention Steering of Large Language Models for Robust Code Understanding under Obfuscation

Aug 2026 · 0 citations · 31 references
Computer Science

TL;DR

This work proposes CodeSteer, a novel attention steering approach that reallocates model attention toward semantically relevant program elements, including backward slices for output prediction and control-flow paths for execution reasoning in large language models.

Abstract

Code obfuscation is widely used in software systems and malware to conceal program logic and hinder analysis, posing significant challenges for both human developers and automated tools. While large language models (LLMs) have shown strong capabilities in code understanding, their robustness to obfuscation remains poorly understood. Our preliminary study shows that LLM performance significantly degrades on obfuscated code, suggesting a reliance on superficial lexical cues rather than deep semantic reasoning. To address this limitation, we propose CodeSteer, a novel attention steering approach that reallocates model attention toward semantically relevant program elements, including backward slices for output prediction and control-flow paths for execution reasoning. Our method integrates lightweight program analysis with inference-time attention steering to guide LLMs toward the core input-to-output dependencies of a program. Experiments across multiple models and datasets demonstrate that CodeSteer significantly improves performance on obfuscated code, often recovering comparable accuracy to the level of unobfuscated programs. We also show CodeSteer's practical utility through a case study on buffer overflow detection, highlighting its potential for malware/vulnerability analysis and reverse engineering of obfuscated code.

View source

Similar papers

#artificial intelligence Preprint Aug 2026

Interpreting and Steering for Safe and Correct Code Generation

DuoSteer is proposed, a double-steering approach that simultaneously applies safety and code-correctness steering to attention heads and outperforms not only other steering variants but also prompting and supervised fine-tuning baselines for inference-time vulnerability reduction.

Hao Yan, Zi-Yu Yao · 0 citations
#small language model Preprint Aug 2026

Vulnerable Code Search: Transferable Attack for Code Language Models

This paper introduces a programming language-agnostic, transferable, adversarial attack that exploits this CLM vulnerability and demonstrates that this attack, even when computed using smaller code embedding models, is highly effective and transferable to larger, closed-source embedding models.

Kaicheng Wang, Liyan Huang, Jesse Thomason et al. · 0 citations
#artificial intelligence Review Sep 2026

Complexity-Aware Evaluation of LLM Comprehension

Large language models (LLMs) are increasingly used for software engineering tasks that require understanding existing source code, including behavior prediction, function explanation, debugging, and code review. However, aggregate benchmark accuracy can conceal how model reliability changes as source code becomes struc...

Ali Mohammadi Esfahani, Nafiseh Kahani, Samuel A.Ajila · 0 citations
Open access Aug 2026

How well do LLMs understand code?

SemBench is introduced, a novel benchmark consisting of 1000 diverse C programs sourced from the CodeParrot GitHub-code dataset, with 15,404 semantic questions spanning six basic but fundamental properties: dead code-statement, data dependency, function reachability, dominator, dead code-loop, and liveness.

Jade Xu, Ren-Liang Sun, Zijian Ding et al. · 0 citations
#small language model Preprint Sep 2026

Towards Behavior Tree-Guided Vulnerability Detection with Lightweight LLMs

Large Language Models (LLMs) are increasingly used for software vulnerability detection, but their performance depends on how source code is represented in the input. Most prompting approaches use source code in its original form, while some works propose the use of structured representations. Abstract Syntax Trees (AS...

Enna Bašić, A. Giaretta · 0 citations
#artificial intelligence Book Open access Sep 2026

Predicting Program Exit Code with LLMs and Programming Language Semantics

This work evaluates open-source coding LLMs under two semantic formalisms and two semantic shifts across Human-Written, LLM-Translated, and Fuzzer-Generated program splits, showing that LLMs lean on pre-training priors rather than systematically applying the given rules.

Lara Marinov, Aditya Thimmaiah, Jayanth Srinivasa et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.