Skip to content
Open access

On Behavioral Alignment of Model-Code and Human-Code Understandability via Behavioral Proxies

Aug 2026 · Proceedings of the ACM on Software Engineering · Vol 3, pp. 2543 - 2565 · 0 citations · 46 references
Computer Science

TL;DR

It is demonstrated that semantic self-consistency is a reliable and extensible measure to be used as a behavioral proxy for quantifying model code understandability, with broad implications in both software engineering research and practice.

Abstract

Code understandability is a critical aspect of software quality. Prior research has largely focused on this attribute from a human-centric or code-centric perspective, while it should be viewed as a relational property arising from the interaction between a reader and the code. With the increasing adoption of large language models (LLMs) in software engineering, we posit that the notion of “reader” should be generalized to encompass both humans and models. Building on this relational perspective, we extend the concept of code understandability to distinguish between human and model code understandability, aiming to investigate the behavioral alignment between the two. To this end, we use a dataset from a prior study containing human-rated judgments of understandability across diverse participant groups, and evaluate multiple open and closed-source LLMs. To operationalize such behavioral alignment, we introduce four behavioral proxies of model-code understandability (BPMU): P0, based on program comprehension-focused question answering; and P1–P3, based on program intent summarization (derived from our proposed semantic self-consistency). Our findings show that LLMs exhibit stronger behavioral alignment with general human code understandability than previously used shallow machine learning baselines. Stratified analyses further reveal that model code understandability aligns most closely with professional developers, but significantly less with other student groups. Role-conditioned prompting does not improve the alignment for the latter, suggesting that current LLMs cannot accurately mimic the perspectives of human readers with varying expertise. Finally, we demonstrate that semantic self-consistency is a reliable and extensible measure to be used as a behavioral proxy for quantifying model code understandability, with broad implications in both software engineering research and practice.

Read PDF

Similar papers

Open access Aug 2026

Software comprehension in code-centric and model-driven settings: an experimental comparison of models and code

It is demonstrated that models and code achieve comparable overall correctness, and thus models alone may be sufficient in model-centric scenarios where access to code is limited or unavailable, and a consistent structure-behavior comprehension gap is revealed.

Iris Reinhartz-Berger, Monique Snoeck · 0 citations

A Facet-Based Perspective on Code 1 Comprehension Task Design for Empirical Research

This work introduces a facet-based perspective that distinguishes between analytical skills, abstraction, and critical evaluation and proposes a method to guide researchers from research goals to concrete tasks and study design, including considerations for question formulation and response formats.

Luisa Bartl, Marvin Wyrich, Sven Apel et al. · 0 citations
Preprint Aug 2026

Evaluating Language Models on Cross-Language Code Functional Equivalence

This work investigates whether LLMs can accurately judge functional equivalence across different programming languages in human-written code, a setting that requires deeper reasoning beyond superficial similarity, and identifies a difficulty-dependent breakdown in equivalence judgment.

Hui Sun, Anderson G. Uchôa, Rohit Gheyi et al. · 0 citations
#natural language process... Preprint Aug 2026

When Who You Are Can Change the Code You Get: A Study of Persona-Induced Bias in LLM Code Generation

Large Language Models (LLMs) are widely used as programming assistants, yet it remains unclear whether and how user's demographic information impacts the technical quality of generated code. We conduct a large-scale empirical study of persona-induced bias in LLM-based code generation, focusing a proprietary model (Gemi...

Anubhav Gupta, M. Figueiredo, L. Machado et al. · 0 citations
Conference Sep 2026

A Model of Software Code Quality Analysis, Assessment, and Improvement System

Existing code analysis systems address individual aspects of software quality (complexity, duplication, style) but do not combine multi-level analysis, metric aggregation, and learning-based inference within a single formal framework, nor do they close the loop between refactoring outcomes and assessment. This paper pr...

Ihor Prokofiev, Oleg Savenko · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.