Skip to content
Book Open access

Can LLMs Really Reason about Code? Studying How Well LLMs Understand the Relation between Input, Code, and Output

Jul 2026 · AIware · pp. 21-30 · 2 citations · 38 references
Computer Science

TL;DR

In this work, the capacity of LLMs to reason about the semantics of code is examined—specifically, their ability to relate code, its inputs, and its outputs to each other, and four datasets are constructed to evaluate diverse program-understanding capabilities.

Abstract

In the past years, large language models (LLMs) have demonstrated remarkable progress in code generation. However, their ability to reason about program behavior remains an open challenge—an ability that is relevant for applications including reverse engineering, debugging, secure code generation, test-driven synthesis, input reconstruction, reverse fuzzing, behavioral monitoring, and safe execution modeling. To study this ability, we examine the capacity of LLMs to reason about the semantics of code—specifically, their ability to relate code, its inputs, and its outputs to each other. To this end, we investigate whether and how well LLMs can predict one of these three components given the other two—that is, (1) predict the input given code and output, (2) predict the output given code and input, and (3) predict the code given input and output. This way, we assess how well LLMs can reason about and understand the underlying relationships that govern program execution. We construct four datasets covering string processing, array operations, and coding challenges in JavaScript and Python to evaluate diverse program-understanding capabilities, incorporating various code mutation techniques to increase complexity. In our evaluation on tasks covering string processing, array operations, and coding challenges, we find that closed-weight models achieve the strongest performance across all datasets, including perfect input recovery on deterministic string tasks. Across tasks, output prediction is comparatively stable, whereas code prediction remains the hardest setting and often fails for smaller models. Finally, cross-codebase transfer is feasible, especially for input prediction, but highly sensitive to model capacity and fine-tuning strategy.

Read PDF

Similar papers

Open access Aug 2026

How well do LLMs understand code?

SemBench is introduced, a novel benchmark consisting of 1000 diverse C programs sourced from the CodeParrot GitHub-code dataset, with 15,404 semantic questions spanning six basic but fundamental properties: dead code-statement, data dependency, function reachability, dominator, dead code-loop, and liveness.

Jade Xu, Ren-Liang Sun, Zijian Ding et al. · 0 citations
Preprint Aug 2026

On the Robustness of LLMs'Internal Representation of Code Correctness

This work studies an internal signal of code correctness that is able to judge candidate solutions better than the model's token-level or stated confidence, leaving open an important question: whether it reflects a robust property of the model or an artifact of that choice.

Francisco Ribeiro, Sohaila Abdulsattar, R. Gonzalez et al. · 1 citation
#artificial intelligence Preprint Aug 2026

Interpreting and Steering for Safe and Correct Code Generation

DuoSteer is proposed, a double-steering approach that simultaneously applies safety and code-correctness steering to attention heads and outperforms not only other steering variants but also prompting and supervised fine-tuning baselines for inference-time vulnerability reduction.

Hao Yan, Zi-Yu Yao · 0 citations
#small language model Preprint Sep 2026

Code Transformation Rule Synthesis using LLMs: Potential and Limits

Due to their black-box nature, LLMs suffer from limited explain- ability and a lack of determinism. Their usage cost can also rise, particularly with repetitive tasks on large codebases. To mitigate this, we conduct a novel empirical study targeting three domain- specific languages for transformation rules, namely Comb...

Axel Allain, Aymeric Blot, D. Khelladi et al. · 1 citation
#artificial intelligence Book Sep 2026

Predicting Program Exit Code with LLMs and Programming Language Semantics

This work evaluates open-source coding LLMs under two semantic formalisms and two semantic shifts across Human-Written, LLM-Translated, and Fuzzer-Generated program splits, showing that LLMs lean on pre-training priors rather than systematically applying the given rules.

Lara Marinov, Aditya Thimmaiah, Jayanth Srinivasa et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.