Skip to content
Preprint

NEUROTESTGEN: Neuro-Symbolic Guided Test Generation with Large Language Models

Sep 2026 · 0 citations · 28 references
Computer Science

TL;DR

This paper introduces NEUROTESTGEN, a hybrid approach that integrates symbolic execution with LLM-driven test synthesis to generate test cases targeting on-demand code coverage and incorporates an iterative feedback loop that validates LLM-generated tests and provides corrective guidance until the target line or branch is covered or a limit is reached.

Abstract

Ensuring high structural coverage remains a fundamental challenge in automated test generation, particularly for complex software systems where reaching specific lines or branches requires satisfying intricate control- and data-flow constraints. Large Language Models (LLMs) have recently demonstrated strong capabilities in producing human-like test cases; however, they often struggle to generate inputs that satisfy precise path conditions. Conversely, symbolic execution can systematically derive such constraints, but it often fails to construct realistic, executable test cases and is constrained by scalability limitations. In this paper, we introduce NEUROTESTGEN, a hybrid approach that integrates symbolic execution with LLM-driven test synthesis to generate test cases targeting on-demand code coverage. Given a set of target statements within a method, NEUROTESTGEN first employs a symbolic analysis engine (i.e., the Z3 SMT solver) to extract path-specific constraints and construct a symbolic guidance specification for the desired coverage goal. This specification is then used to guide an LLM in synthesizing concrete test cases that are both structurally valid and semantically meaningful. For paths involving complex object-related constraints that are difficult for SMT solvers to handle, NEUROTESTGEN leverages LLMs to infer plausible constraints. Furthermore, NEUROTESTGEN incorporates an iterative feedback loop that validates LLM-generated tests and provides corrective guidance until the target line or branch is covered or a limit is reached. Our empirical evaluation on a widely used benchmark demonstrates that NEUROTESTGEN significantly outperforms the state-of-the-art approach across multiple LLMs, including Llama 3.3 70B1, GPT-4o Mini, Claude 3.5 Haiku3, and Claude Sonnet 4.6.

View source

Similar papers

Preprint Sep 2026

NeuroSTAR: Automata-guided Neuro-symbolic Specification Formalization

Automated translation of natural language (NL) descriptions into Linear Temporal Logic over finite traces (LTLf) is a prerequisite for automated formal verification of a system's dynamic behavior. Several LLM-based methods have recently shown potential for this task. However, they struggle with the nuance of natural la...

Joy Saha, Trey Woodlief, Sebastian G. Elbaum et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Coverage-Driven RTL Assertion Generation with Formal Exploration and Neuro-Symbolic Refinement

NeuroAssertion is presented, a coverage-driven assertion generation framework that combines formal trace generation, syntax-guided synthesis (SyGuS), and an agent-inspired refinement process within a unified framework that delivers around 2X more assertions and about 2X higher mutation coverage than traditional asserti...

Zhi-Yuan Yan, Ziyue Zheng, Hongce Zhang · 0 citations
Preprint Aug 2026

Neuro-Symbolic Proof-of-Vulnerability Generation with Open-Weight Models

POVGEN, a low-cost neuro-symbolic framework that makes PoV generation cost-effective via semantic focusing and LLM-guided constraint reasoning using open-weight models, and fine-tuned open-weight models match frontier commercial LLMs on key sub-tasks while running locally at no per-sample API cost.

Yu Nong, Hai-Peng Cai · 0 citations
#software testing Preprint Aug 2026

POLYFLOW: A Neuro-Symbolic Framework for Static Cross-Language Information Flow Analysis

PolyFlow is a neural-symbolic framework for statically reasoning about information flow across language boundaries, combining large language models (LLMs) and static analysis synergistically, and is cost-effective and superior to various kinds of state-of-the-art baselines.

Haoran Yang, Zhi-Xuan Zhong, Jiawei Guo et al. · 0 citations
#software testing Preprint Sep 2026

Enhancing Automated Unit Test Generation for NLP Libraries Using Large Language Models

LLMSuite is proposed, a hybrid test generation framework that integrates self-refinement prompting with class-level LLM reasoning into the search-based testing process and complements manually written test suites by exercising domain-specific behaviors that are often left untested.

Amirhossein Deljouyi, Annibale Panichella, Andy Zaidman · 0 citations
Book Open access Aug 2026

SymDiag: Explainable Diagnosis for LLM Reasoning via Neuro-Symbolic Verification

Large language models (LLMs) increasingly serve as data-driven reasoners, yet their chains-of-thought (CoT) can be unfaithful even when final answers are correct. Most existing ''verification'' signals are not diagnostic: answer matching observes only the outcome, LLM-as-judge provides subjective and non-verifiable cri...

Wenyao Cui, Hua-Ping Zhang, Yongyi Huang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.