Jul 2026· AIware· pp. 21-30· 2 citations· 38 references
Computer Science
TL;DR
In this work, the capacity of LLMs to reason about the semantics of code is examined—specifically, their ability to relate code, its inputs, and its outputs to each other, and four datasets are constructed to evaluate diverse program-understanding capabilities.
Abstract
In the past years, large language models (LLMs) have demonstrated remarkable progress in code generation. However, their ability to reason about program behavior remains an open challenge—an ability that is relevant for applications including reverse engineering, debugging, secure code generation, test-driven synthesis, input reconstruction, reverse fuzzing, behavioral monitoring, and safe execution modeling. To study this ability, we examine the capacity of LLMs to reason about the semantics of code—specifically, their ability to relate code, its inputs, and its outputs to each other. To this end, we investigate whether and how well LLMs can predict one of these three components given the other two—that is, (1) predict the input given code and output, (2) predict the output given code and input, and (3) predict the code given input and output. This way, we assess how well LLMs can reason about and understand the underlying relationships that govern program execution. We construct four datasets covering string processing, array operations, and coding challenges in JavaScript and Python to evaluate diverse program-understanding capabilities, incorporating various code mutation techniques to increase complexity. In our evaluation on tasks covering string processing, array operations, and coding challenges, we find that closed-weight models achieve the strongest performance across all datasets, including perfect input recovery on deterministic string tasks. Across tasks, output prediction is comparatively stable, whereas code prediction remains the hardest setting and often fails for smaller models. Finally, cross-codebase transfer is feasible, especially for input prediction, but highly sensitive to model capacity and fine-tuning strategy.
SemBench is introduced, a novel benchmark consisting of 1000 diverse C programs sourced from the CodeParrot GitHub-code dataset, with 15,404 semantic questions spanning six basic but fundamental properties: dead code-statement, data dependency, function reachability, dominator, dead code-loop, and liveness.
Jade Xu, Ren-Liang Sun, Zijian Ding et al.· Communications AI & Computin...· 0 citations
This work studies an internal signal of code correctness that is able to judge candidate solutions better than the model's token-level or stated confidence, leaving open an important question: whether it reflects a robust property of the model or an artifact of that choice.
Francisco Ribeiro, Sohaila Abdulsattar, R. Gonzalez et al.· 1 citation
DuoSteer is proposed, a double-steering approach that simultaneously applies safety and code-correctness steering to attention heads and outperforms not only other steering variants but also prompting and supervised fine-tuning baselines for inference-time vulnerability reduction.
Due to their black-box nature, LLMs suffer from limited explain- ability and a lack of determinism. Their usage cost can also rise, particularly with repetitive tasks on large codebases. To mitigate this, we conduct a novel empirical study targeting three domain- specific languages for transformation rules, namely Comb...
Axel Allain, Aymeric Blot, D. Khelladi et al.· 1 citation
This work evaluates open-source coding LLMs under two semantic formalisms and two semantic shifts across Human-Written, LLM-Translated, and Fuzzer-Generated program splits, showing that LLMs lean on pre-training priors rather than systematically applying the given rules.
Lara Marinov, Aditya Thimmaiah, Jayanth Srinivasa et al.· 0 citations
The results show that executable feedback can repair secure-code generation, but its benefits depend on the model, task, feedback entry point, and especially test coverage.
Yun-Hao Liang, Cheng-Guang Gan, Rui-Xuan Ying et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.