Aug 2026· Proceedings of the ACM on Software Engineering· Vol 3, pp. 2543 - 2565· 0 citations· 46 references
Computer Science
TL;DR
It is demonstrated that semantic self-consistency is a reliable and extensible measure to be used as a behavioral proxy for quantifying model code understandability, with broad implications in both software engineering research and practice.
Abstract
Code understandability is a critical aspect of software quality. Prior research has largely focused on this attribute from a human-centric or code-centric perspective, while it should be viewed as a relational property arising from the interaction between a reader and the code. With the increasing adoption of large language models (LLMs) in software engineering, we posit that the notion of “reader” should be generalized to encompass both humans and models. Building on this relational perspective, we extend the concept of code understandability to distinguish between human and model code understandability, aiming to investigate the behavioral alignment between the two. To this end, we use a dataset from a prior study containing human-rated judgments of understandability across diverse participant groups, and evaluate multiple open and closed-source LLMs. To operationalize such behavioral alignment, we introduce four behavioral proxies of model-code understandability (BPMU): P0, based on program comprehension-focused question answering; and P1–P3, based on program intent summarization (derived from our proposed semantic self-consistency). Our findings show that LLMs exhibit stronger behavioral alignment with general human code understandability than previously used shallow machine learning baselines. Stratified analyses further reveal that model code understandability aligns most closely with professional developers, but significantly less with other student groups. Role-conditioned prompting does not improve the alignment for the latter, suggesting that current LLMs cannot accurately mimic the perspectives of human readers with varying expertise. Finally, we demonstrate that semantic self-consistency is a reliable and extensible measure to be used as a behavioral proxy for quantifying model code understandability, with broad implications in both software engineering research and practice.
It is demonstrated that models and code achieve comparable overall correctness, and thus models alone may be sufficient in model-centric scenarios where access to code is limited or unavailable, and a consistent structure-behavior comprehension gap is revealed.
Iris Reinhartz-Berger, Monique Snoeck· Journal of Software and Syst...· 0 citations
This work introduces a facet-based perspective that distinguishes between analytical skills, abstraction, and critical evaluation and proposes a method to guide researchers from research goals to concrete tasks and study design, including considerations for question formulation and response formats.
Luisa Bartl, Marvin Wyrich, Sven Apel et al.· 0 citations
This work investigates whether LLMs can accurately judge functional equivalence across different programming languages in human-written code, a setting that requires deeper reasoning beyond superficial similarity, and identifies a difficulty-dependent breakdown in equivalence judgment.
Hui Sun, Anderson G. Uchôa, Rohit Gheyi et al.· 0 citations
These findings position inline comments as model-sensitive latent semantic prompts, with implications for AI-in-the-loop development and design of comment conventions for AI-assisted maintenance.
Angela N. Johnson· Frontiers in Artificial Inte...· 0 citations
Large Language Models (LLMs) are widely used as programming assistants, yet it remains unclear whether and how user's demographic information impacts the technical quality of generated code. We conduct a large-scale empirical study of persona-induced bias in LLM-based code generation, focusing a proprietary model (Gemi...
Anubhav Gupta, M. Figueiredo, L. Machado et al.· 0 citations
Existing code analysis systems address individual aspects of software quality (complexity, duplication, style) but do not combine multi-level analysis, metric aggregation, and learning-based inference within a single formal framework, nor do they close the loop between refactoring outcomes and assessment. This paper pr...
Ihor Prokofiev, Oleg Savenko· Automation, Control, and Inf...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.