GraphDroid is proposed, an intent-driven GUI testing framework that integrates a cluster-based memory mechanism to effectively identify uncovered functionalities from historically visited states for comprehensive application testing and adopts an asynchronous intent generation paradigm that eliminates the latency bottleneck.
Abstract
Automated GUI testing is a widely adopted technique for ensuring mobile application quality by simulating user interactions to exercise functionalities. Despite the research breakthroughs in the past decades, covering complex functionalities that require multi-step action sequences still remains challenging. Traditional tools lack semantic understanding capability and can rarely synthesize such action sequences. Recent LLM-based tools can generate test intents describing target functionalities and leverage the LLM to fulfill the intents, but suffer from three key limitations: 1) loss of historical context for identifying uncovered functionalities, 2) synchronous intent generation that blocks exploration, and 3) per-step LLM-driven fulfillment incurring high cost and latency. To address these limitations, we propose GraphDroid, an intent-driven GUI testing framework that integrates a cluster-based memory mechanism to effectively identify uncovered functionalities from historically visited states for comprehensive application testing. For improving testing efficiency, GraphDroid adopts an asynchronous intent generation paradigm that eliminates the latency bottleneck and a hybrid intent fulfillment strategy that reserves the LLM for fulfilling complex intents while delegating simple intents to a lightweight heuristic algorithm. We evaluate GraphDroid on 41 real-world Android apps against six state-of-the-art baselines. Results show that GraphDroid outperforms all baselines, achieving up to 36.4% higher code coverage while incurring less than one eighth of the cost of the best pure LLM-based baseline. GraphDroid also exposes 19 bugs in the 41 apps and detects 13 of 52 crashes in the Themis bug benchmark, surpassing all the six baselines. Seven of the 19 bugs were previously unknown and we reported them to the developers. So far, four bugs have been confirmed and fixed.
The commercialization of the mobile Internet has driven many apps to adopt deceptive interaction designs (referred to as interaction flow contamination), which misalign system behaviors with user true intent and may harm user interests. Existing detection approaches suffer from semantic comprehension gaps, insufficient...
Jia-Lu Li, Xiang Li· Proceedings of the 2026 ACM...· 0 citations
Large language models (LLMs) are increasingly used in software pipelines, raising concerns about harmful behaviors in security-critical domains. Existing safety evaluations predominantly probe models with single prompts or short interactions, and therefore do not capture how safety behaves under multi-step workflows wh...
Lu Yan, Zhuo Zhang, Xiang-Zhe Xu et al.· Proceedings of the ACM on So...· 0 citations
As Android applications grow in scale, latent performance issues increasingly degrade user experience and business outcomes, yet systematically identifying optimization opportunities in large codebases remains challenging. We present OptiMine, a hybrid knowledge-to-code framework that integrates large language models (...
Peng-Bo Du, Qiu-Ping Yi, Liang-Zheng Zhang et al.· Proceedings of the ACM on So...· 0 citations
Graphical user interface (GUI) tests are widely used in regression testing of mobile applications (apps) to validate app behavior from the user's perspective. However, frequent app evolution, such as UI redesigns and feature updates, often renders existing GUI tests obsolete, even when underlying functionality remains...
Shi-Wen Song, Yiheng Xiong, Wen-Bo Guo et al.· Proceedings of the ACM on So...· 0 citations
The paper asks whether high-level conceptual specifications improve LLM-generated PBT quality and whether they help developers extend and maintain AI-generated systems (RQ2), and reports preliminary results applying Spinach to two open-source applications.
Savitha Ravi, Michael Coblenz· Proceedings of the 1st Inter...· 0 citations
ProgramDistill is introduced, a benchmark evaluating coding agents on features discovered through interaction with fully functional reference applications, and provides a scalable benchmark with controlled difficulty for evaluating and diagnosing coding agents, and a natural basis for future curriculum-based training.
Jeonghye Kim, Minseon Kim, Young Jin Kim et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.