Skip to content
Preprint

GraphDroid: Asynchronous LLM-Based Mobile App GUI Testing via History-Aware Exploration and Hybrid Intent Fulfillment

Sep 2026 · 0 citations · 50 references
Computer Science

TL;DR

GraphDroid is proposed, an intent-driven GUI testing framework that integrates a cluster-based memory mechanism to effectively identify uncovered functionalities from historically visited states for comprehensive application testing and adopts an asynchronous intent generation paradigm that eliminates the latency bottleneck.

Abstract

Automated GUI testing is a widely adopted technique for ensuring mobile application quality by simulating user interactions to exercise functionalities. Despite the research breakthroughs in the past decades, covering complex functionalities that require multi-step action sequences still remains challenging. Traditional tools lack semantic understanding capability and can rarely synthesize such action sequences. Recent LLM-based tools can generate test intents describing target functionalities and leverage the LLM to fulfill the intents, but suffer from three key limitations: 1) loss of historical context for identifying uncovered functionalities, 2) synchronous intent generation that blocks exploration, and 3) per-step LLM-driven fulfillment incurring high cost and latency. To address these limitations, we propose GraphDroid, an intent-driven GUI testing framework that integrates a cluster-based memory mechanism to effectively identify uncovered functionalities from historically visited states for comprehensive application testing. For improving testing efficiency, GraphDroid adopts an asynchronous intent generation paradigm that eliminates the latency bottleneck and a hybrid intent fulfillment strategy that reserves the LLM for fulfilling complex intents while delegating simple intents to a lightweight heuristic algorithm. We evaluate GraphDroid on 41 real-world Android apps against six state-of-the-art baselines. Results show that GraphDroid outperforms all baselines, achieving up to 36.4% higher code coverage while incurring less than one eighth of the cost of the best pure LLM-based baseline. GraphDroid also exposes 19 bugs in the 41 apps and detects 13 of 52 crashes in the Themis bug benchmark, surpassing all the six baselines. Seven of the 19 bugs were previously unknown and we reported them to the developers. So far, four bugs have been confirmed and fixed.

View source

Similar papers

Book Open access Oct 2026

Poster: Detecting Intent Inconsistency in Mobile Apps:A Focus on Interaction Flow Contamination

The commercialization of the mobile Internet has driven many apps to adopt deceptive interaction designs (referred to as interaction flow contamination), which misalign system behaviors with user true intent and may harm user interests. Existing detection approaches suffer from semantic comprehension gaps, insufficient...

Jia-Lu Li, Xiang Li · 0 citations
Open access Oct 2026

CAST: A Compiler-Based Framework for Systematically Testing LLM Compositional Safety

Large language models (LLMs) are increasingly used in software pipelines, raising concerns about harmful behaviors in security-critical domains. Existing safety evaluations predominantly probe models with single prompts or short interactions, and therefore do not capture how safety behaves under multi-step workflows wh...

Lu Yan, Zhuo Zhang, Xiang-Zhe Xu et al. · 0 citations
Open access Oct 2026

OptiMine: Scalable and Precise Code Optimization for Android Apps via LLM-Driven Semantic Analysis

As Android applications grow in scale, latent performance issues increasingly degrade user experience and business outcomes, yet systematically identifying optimization opportunities in large codebases remains challenging. We present OptiMine, a hybrid knowledge-to-code framework that integrates large language models (...

Peng-Bo Du, Qiu-Ping Yi, Liang-Zheng Zhang et al. · 0 citations
Open access Oct 2026

Characterizing and Repairing Obsolete Android GUI Tests under UI Evolution

Graphical user interface (GUI) tests are widely used in regression testing of mobile applications (apps) to validate app behavior from the user's perspective. However, frequent app evolution, such as UI redesigns and feature updates, often renders existing GUI tests obsolete, even when underlying functionality remains...

Shi-Wen Song, Yiheng Xiong, Wen-Bo Guo et al. · 0 citations
#software testing Book Open access Oct 2026

SPINACH: Inferring Properties of Web Applications for Property-Based Testing

The paper asks whether high-level conceptual specifications improve LLM-generated PBT quality and whether they help developers extend and maintain AI-generated systems (RQ2), and reports preliminary results applying Spinach to two open-source applications.

Savitha Ravi, Michael Coblenz · 0 citations
#artificial intelligence Preprint Sep 2026

ProgramDistill: From Interactive Web Apps to Verifiable Reference-Guided SWE Tasks

ProgramDistill is introduced, a benchmark evaluating coding agents on features discovered through interaction with fully functional reference applications, and provides a scalable benchmark with controlled difficulty for evaluating and diagnosing coding agents, and a natural basis for future curriculum-based training.

Jeonghye Kim, Minseon Kim, Young Jin Kim et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.