A Multilayered Evaluation Framework is introduced that decouples deterministic database logic from flexible AI semantics and achieves state-of-the-art overall accuracy across both platforms, proposing a reliable standard for benchmarking AI-powered SQL generators.
Abstract
SQL has been augmented with AI operators, enabling modern data analytics platforms to derive insights from both structured and unstructured data. We observe that while current Text-to-SQL systems can successfully generate these AI-augmented queries, reliably evaluating their correctness remains a critical open challenge. Current metrics, which rely on exact query results and deterministic execution, systematically fail against the flexible, non-deterministic outputs of AI operators. In this paper, we formalize these unique evaluation failure modes and introduce a Multilayered Evaluation Framework that decouples deterministic database logic from flexible AI semantics. We test our approach across both industry (BigQuery) and academic (ThalamusDB) systems. We demonstrate that traditional Execution Accuracy severely penalizes valid queries, achieving as low as a 25% detection rate for correct translations. Furthermore, even a state-of-the-art LLM-based autorater falsely rejects 32% of accurate queries due to the complexity of judging both relational and AI components simultaneously. By validating the standard relational logic and the AI operations separately, our framework achieves state-of-the-art overall accuracy across both platforms (up to 97.2%), proposing a reliable standard for benchmarking AI-powered SQL generators.
This work formalizes Text-to-SQL verification as a standalone task: assessing whether a candidate SQL query semantically satisfies a natural language request, without access to ground-truth labels, and proposes and compares two modular verification strategies.
Tarfah Alrashed, Madhup Sukoon, David R Karger et al.· Proceedings of the VLDB Endo...· 0 citations
Large language models have been increasingly applied for Text-to-SQL tasks recently, due to the fact that generating correct SQL queries is the most important factor. This study presents a comparison of AI agents based on open-source and closed-source large language models for generating SQL queries. GPT-4o, Claude 3.7...
Julia Sierpień, M. Skublewska-Paszkowska· Journal of Computer Sciences...· 0 citations
This work shows that introspection and sampling are complementary but not disjoint on BIRD-Interact-Lite, a Text-to-SQL benchmark with 300 tasks, annotated ambiguities, and an LLM-based user simulator, and proposes a proposed multi-label grammar that substantially narrows it without affecting downstream context.
Text-to-SQL models translate natural language questions into SQL, enabling non-technical users to access databases. However, most existing research focuses on correctness, neglecting query efficiency. In this paper, we address the challenge of evaluating the execution efficiency of generated SQL in Text-to-SQL by intro...
Li Lin, Yun-Feng Shen, Ling-Feng Bao et al.· Proceedings of the ACM on So...· 0 citations
LIMIT(Less Is More for Instruction Tuning in Text-to-SQL), a data-centric framework that demonstrates strong database reasoning can emerge from an extremely compact training set when examples are strategically selected, is proposed, suggesting that careful data curation, rather than scale, is the key to efficient Text-...
Hao-Yuan Ma, Heng-Wei Liu, Linjuan Wu et al.· 0 citations
BIRD-History is introduced, a benchmark consisting of 1,393 tasks across 11 databases, designed to evaluate text-to-SQL systems'ability to ground underspecified natural language questions using historical SQL scripts, and a plug-in retriever that extracts five types of external knowledge from historical SQL scripts, th...
Yun-Fan Zhou, Qi-Ming Shi, Yi-Zhou Yang et al.· 0 citations
With $2.1 million funding from Google.org, the open-source Public Transit Intelligence Hub will unify public transit monitoring, operations, and passenger communication.
Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.