Skip to content

The Stochastic Shift: A New Evaluation Paradigm for Text-to-SQL with AI Operators

Sep 2026 · 0 citations · 17 references
Computer Science

TL;DR

A Multilayered Evaluation Framework is introduced that decouples deterministic database logic from flexible AI semantics and achieves state-of-the-art overall accuracy across both platforms, proposing a reliable standard for benchmarking AI-powered SQL generators.

Abstract

SQL has been augmented with AI operators, enabling modern data analytics platforms to derive insights from both structured and unstructured data. We observe that while current Text-to-SQL systems can successfully generate these AI-augmented queries, reliably evaluating their correctness remains a critical open challenge. Current metrics, which rely on exact query results and deterministic execution, systematically fail against the flexible, non-deterministic outputs of AI operators. In this paper, we formalize these unique evaluation failure modes and introduce a Multilayered Evaluation Framework that decouples deterministic database logic from flexible AI semantics. We test our approach across both industry (BigQuery) and academic (ThalamusDB) systems. We demonstrate that traditional Execution Accuracy severely penalizes valid queries, achieving as low as a 25% detection rate for correct translations. Furthermore, even a state-of-the-art LLM-based autorater falsely rejects 32% of accurate queries due to the complexity of judging both relational and AI components simultaneously. By validating the standard relational logic and the AI operations separately, our framework achieves state-of-the-art overall accuracy across both platforms (up to 97.2%), proposing a reliable standard for benchmarking AI-powered SQL generators.

View source

Similar papers

Jul 2026

Developing and Benchmarking Verification Algorithms to Improve Text-to-SQL Generation

This work formalizes Text-to-SQL verification as a standalone task: assessing whether a candidate SQL query semantically satisfies a natural language request, without access to ground-truth labels, and proposes and compares two modular verification strategies.

Tarfah Alrashed, Madhup Sukoon, David R Karger et al. · 0 citations
Open access Sep 2026

Comparison of AI agents for creating SQL queries

Large language models have been increasingly applied for Text-to-SQL tasks recently, due to the fact that generating correct SQL queries is the most important factor. This study presents a comparison of AI agents based on open-source and closed-source large language models for generating SQL queries. GPT-4o, Claude 3.7...

Julia Sierpień, M. Skublewska-Paszkowska · 0 citations

Two Surfaces of Ambiguity: Complementary Detection for Text-to-SQL

This work shows that introspection and sampling are complementary but not disjoint on BIRD-Interact-Lite, a Text-to-SQL benchmark with 300 tasks, annotated ambiguities, and an LLM-based user simulator, and proposes a proposed multi-label grammar that substantially narrows it without affecting downstream context.

Leonhard Liu, Patrick K. Erdelt, Two · 1 citation
Open access Oct 2026

Cracking Query Bottlenecks: Towards Efficiency-Oriented Text-to-SQL Generation

Text-to-SQL models translate natural language questions into SQL, enabling non-technical users to access databases. However, most existing research focuses on correctness, neglecting query efficiency. In this paper, we address the challenge of evaluating the execution efficiency of generated SQL in Text-to-SQL by intro...

Li Lin, Yun-Feng Shen, Ling-Feng Bao et al. · 0 citations
#artificial intelligence Preprint Sep 2026

LIMIT: Less Is More for Instruction Tuning in Text-to-SQL

LIMIT(Less Is More for Instruction Tuning in Text-to-SQL), a data-centric framework that demonstrates strong database reasoning can emerge from an extremely compact training set when examples are strategically selected, is proposed, suggesting that careful data curation, rather than scale, is the key to efficient Text-...

Hao-Yuan Ma, Heng-Wei Liu, Linjuan Wu et al. · 0 citations
#artificial intelligence Preprint Aug 2026

BIRD-History: A Benchmark for History-Driven Text-to-SQL with Fine-Grained Knowledge Annotations

BIRD-History is introduced, a benchmark consisting of 1,393 tasks across 11 databases, designed to evaluate text-to-SQL systems'ability to ground underspecified natural language questions using historical SQL scripts, and a plug-in retriever that extracts five types of external knowledge from historical SQL scripts, th...

Yun-Fan Zhou, Qi-Ming Shi, Yi-Zhou Yang et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.