Skip to content

DiffTestGen: Change-Directed LLM-Based Testing for Exposing Behavioral Differences

Jul 2026 · arXiv.org · Vol abs/2607.16024 · 0 citations · 43 references
Computer Science

TL;DR

DiffTestGen is presented, a novel change-directed, LLM-based differential testing approach specifically designed to expose behavioral differences introduced by a code change and it is shown that the identified behavioral differences can be used to detect regression bugs missed by the best existing approaches.

Abstract

As software evolves over time, it is important to ensure that any behavioral changes occur as intended by developers. A promising approach for this goal is to generate tests that expose behavioral differences between the old and new versions of a program. However, current approaches fail to trigger behavioral differences for many code changes. This paper presents~DiffTestGen, a novel change-directed, LLM-based differential testing approach specifically designed to expose behavioral differences introduced by a code change. The approach is enabled by two key contributions: First, DiffTestGen leverages static call graph analysis and project documentation to identify valid entry points for test generation and to guide the LLM toward reaching the changed code. Second, DiffTestGen iteratively improves our newly introduced union coverage metric, which combines coverage of modified code in the old and the new version, by providing targeted coverage feedback to the LLM. We evaluate DiffTestGen on two datasets comprising a total of 463 PRs. DiffTestGen exposes behavioral differences in 78.2% of the PRs while achieving an average union coverage of 90.7%. Compared with the baselines, DiffTestGen exposes 99 more PRs overall and increases code coverage by 12.5% and 15.6% percentage points, respectively. By integrating DiffTestGen with the Testora regression detector, we show that the identified behavioral differences can be used to detect regression bugs missed by the best existing approaches.

View source

Similar papers

Preprint Aug 2026

Detecting Behavioral Changes in Python Refactoring Implementations with Foundation Models

This work proposes an approach based on a foundation model oracle that analyzes git-style diffs to identify behavioral changes introduced by Python refactorings and uncovered 13 distinct bugs among the seven refactoring types studied.

Jonhnanthan Oliveira, Rohit Gheyi, Márcio Ribeiro et al. · 0 citations
Jul 2026

Evaluating and Mitigating the Misguidance Effect of Buggy Code in LLM-Generated Unit Tests

A specification-based unit test generation paradigm that replaces the code under test in the prompt with an LLM-generated specification docstring is introduced, showing that this paradigm effectively reduces misguided tests while substantially increasing effective tests, improves multi-round, feedback-driven test generation pipelines, and remains applicable to both buggy and bug-free code.

Junda Zhao, Shurui Zhou, Eldan Cohen · 0 citations
#software testing Preprint Aug 2026

BreakGuard: Towards Detecting Dependency Breaking Changes with LLM-Generated Tests

This work proposes BreakGuard, an approach that generates a test suite to detect breaking changes in clients and successfully detected BCs from different library categories, but finds LLM-generated tests to be more reliable for detecting crash-type breaking changes as opposed to behavioural BCs.

Rachna Raj, Benoit Baudry, Diego Elias Costa · 0 citations
Jul 2026

Do Coverage and Mutation Scores of LLM-Generated Test Suites Correlate with Their Effectiveness? (Replicability Study)

It is found that the usefulness of coverage and mutation is highly context-dependent: in regression-style settings where the code provided to the LLM can be reasonably assumed bug-free, these metrics can provide meaningful signals when comparing across models; in another common scenario where the code-under-test may already be buggy and the goal is to expose the bug within the code-under-test, they no longer serve as reliable indicators.

Junda Zhao, Shurui Zhou, Eldan Cohen · 0 citations
Open access Aug 2026

Improving Bug Detection in LLM-Generated Unit Tests: Revisiting Test-Oracle Reliability Across Modern Large Language Models

This paper presents a formal mathematical model for categorizing the outcome of generated-tests into four classes, a couple of basic metrics: Bug-Revealing Rate (BRR) and Bug-Validating Rate (BVR); and two basic statistical tests to ensure that the results are rigorous.

Zeyad Farooq Lutfi · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.