Skip to content

TQTS-Bench: A Multi-Syntax Benchmark for Text-to-Query over Time-Series Databases

Sep 2026 · 0 citations · 25 references
Computer Science

TL;DR

TQTS-BENCH is introduced, a multi-syntax benchmark for evaluating text-to-query capabilities over TSDBs, and error analysis reveals that this performance gap mainly stems from the heterogeneous query syntaxes across different TSDBs, misinterpretation of time-specific intents, and incorrect schema linking.

Abstract

Large language models (LLMs) have significantly advanced natural language querying over relational databases, yet their ability to query time-series databases (TSDBs) remains largely unassessed. Existing benchmarks fail to adequately capture the non-unified query syntaxes, diverse application domains, and unique time-specific query intents inherent to TSDBs. To address this gap, we introduce TQTS-BENCH, a multi-syntax benchmark for evaluating text-to-query capabilities over TSDBs. TQTS-BENCH contains 6,125 high-quality question-answering (QA) pairs spanning 97 TSDBs, 23 distinct query syntaxes, 22 application domains, and 4 types of time-specific query intents. It is constructed through a human-centric AI-assisted workflow, where all QA pairs are carefully reviewed and revised by domain experts to ensure quality and correctness. Extensive evaluations of advanced LLMs and state-of-the-art text-to-query methods reveal challenges in querying TSDBs. Even the best-performing model evaluated, Claude-Opus-5, achieves only 48.98% execution accuracy, while humans reach 87.34%. Error analysis reveals that this performance gap mainly stems from the heterogeneous query syntaxes across different TSDBs, misinterpretation of time-specific intents, and incorrect schema linking. These findings highlight new opportunities to narrow the gap between current LLM capabilities and the requirements of TSDB queries in real-world applications. The benchmark is available at: https://anonymous.4open.science/r/TQTS-Bench-00CD.

View source

Similar papers

#natural language process... Preprint Oct 2026

AptMQL-Bench: From Text-to-SQL to Text-to-MQL via Access-Pattern Schema Design and Data-Preserving Migration

Document databases such as MongoDB are core infrastructure for modern applications, and natural-language interfaces to them---text-to-MQL---would let non-experts query complex, semi-structured data without mastering the query language. Progress on this task depends on high-quality benchmarks, which are most practically...

Hy Nguyen, Nabi Rezvani, Robin Vujanic · 0 citations
Open access Oct 2026

Cracking Query Bottlenecks: Towards Efficiency-Oriented Text-to-SQL Generation

Text-to-SQL models translate natural language questions into SQL, enabling non-technical users to access databases. However, most existing research focuses on correctness, neglecting query efficiency. In this paper, we address the challenge of evaluating the execution efficiency of generated SQL in Text-to-SQL by intro...

Li Lin, Yun-Feng Shen, Ling-Feng Bao et al. · 0 citations
#artificial intelligence Preprint Aug 2026

BIRD-History: A Benchmark for History-Driven Text-to-SQL with Fine-Grained Knowledge Annotations

BIRD-History is introduced, a benchmark consisting of 1,393 tasks across 11 databases, designed to evaluate text-to-SQL systems'ability to ground underspecified natural language questions using historical SQL scripts, and a plug-in retriever that extracts five types of external knowledge from historical SQL scripts, th...

Yun-Fan Zhou, Qi-Ming Shi, Yi-Zhou Yang et al. · 0 citations

Technical Report - A Comparative Evaluation of Schema Subsetting for LLM-based NL-to-SQL over Large-Schema Databases

This paper systematically evaluates 7 real-world schema subsetting modules across 3 contemporary NL-to-SQL benchmarks, and introduces BigBird–an expansion of the Bird benchmark datasets that provides additional data for evaluating subsetting of large schemas, and introduces new subsetting-specific performance and effic...

Kyle Luoma, Arun Kumar · 0 citations
Preprint Aug 2026

MDB-Link: Hierarchical Schema Linking for Multi-Database Text-to-SQL

This work proposes MDB-Link, a hierarchical schema-linking framework that retrieves question-relevant columns from a global index, aggregates retrieval evidence to shortlist databases, and uses a budget-aware large language model (LLM) for database reranking, table selection, and column grounding.

Beiyu Xu, Zhenyu Wu, Jiao-Yan Chen et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.