Skip to content

Policy-Conditioned Constrained Decoding for Column-Level Access Control in Text-to-SQL

Jul 2026 · arXiv.org · Vol abs/2607.12341 · 0 citations · 29 references
Computer Science

TL;DR

A per-token logits mask is applied that deterministically eliminates single-query column-use violations on the supported SQL fragment in a single decoding pass and achieves 0% Leakage Rate and Coverage up to 88.7% on Spider-CU, while staying within +10% tokens of direct prompting.

Abstract

Text-to-SQL is increasingly deployed across trust boundaries between data providers and users. Such deployment must balance three competing requirements: policy compliance, answer coverage, and bounded cost. Existing approaches typically decide refusal based on which columns a query mentions and enforce it stochastically. Whether a query is compliant, however, depends not only on which columns appear but on how they are used, and stochastic enforcement cannot deterministically rule out violations. We formalize this requirement as a column-use policy over semantic use: output, filter condition, and aggregation argument. We integrate the policy by aligning each role with grammar productions tracked by the decoder. The resulting system, PCC-SQL, applies a per-token logits mask that deterministically eliminates single-query column-use violations on the supported SQL fragment in a single decoding pass. Across three benchmarks and three open-source models, PCC-SQL achieves 0% Leakage Rate and Coverage up to 88.7% on Spider-CU, while staying within +10% tokens of direct prompting. We additionally assess semantic alignment with execution accuracy.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

LIMIT: Less Is More for Instruction Tuning in Text-to-SQL

LIMIT(Less Is More for Instruction Tuning in Text-to-SQL), a data-centric framework that demonstrates strong database reasoning can emerge from an extremely compact training set when examples are strategically selected, is proposed, suggesting that careful data curation, rather than scale, is the key to efficient Text-...

Hao-Yuan Ma, Heng-Wei Liu, Linjuan Wu et al. · 0 citations
Preprint Aug 2026

TraceSQL: Traceable Answerability Estimation for Reference-Free Text-to-SQL Verification

Results show that SQL verification can be performed with a lightweight learned model while retaining feature-level evidence for inspecting and diagnosing its predictions, and feature attribution shows that the model relies on both semantic grounding and deterministic SQL-structure signals.

N. Shukla, Debasmita Panda, Srutanik Bhaduri et al. · 0 citations
#artificial intelligence Preprint Sep 2026

BudgetSchemaBench: A Budget-Swept Diagnostic for Schema Context in Text-to-SQL

Data agents over structured sources must fit database schema into the model's context window. Large catalogs can span many databases and thousands of columns, so cost constraints may require choosing between table coverage and serialization detail well before the context window is full. We introduce BudgetSchemaBench,...

Chen Shen · 0 citations

Two Surfaces of Ambiguity: Complementary Detection for Text-to-SQL

This work shows that introspection and sampling are complementary but not disjoint on BIRD-Interact-Lite, a Text-to-SQL benchmark with 300 tasks, annotated ambiguities, and an LLM-based user simulator, and proposes a proposed multi-label grammar that substantially narrows it without affecting downstream context.

Leonhard Liu, Patrick K. Erdelt, Two · 1 citation
#artificial intelligence Preprint Sep 2026

The Stochastic Shift: A New Evaluation Paradigm for Text-to-SQL with AI Operators

A Multilayered Evaluation Framework is introduced that decouples deterministic database logic from flexible AI semantics and achieves state-of-the-art overall accuracy across both platforms, proposing a reliable standard for benchmarking AI-powered SQL generators.

Tarfah Alrashed, Fatma Ozcan, Per Jacobsson et al. · 0 citations
Preprint Aug 2026

SPOC-SQL: Stage-wise Preference Optimization for Controllable Text-to-SQL

SPOC-SQL is proposed, which decomposes Text-to-SQL into four sequential subtasks following standard SQL execution logic and designs stage-specific optimization strategies for the model to learn key decisions, with the objective of enhancing structured decision-making during query construction.

Yingnan Chen, Chun Ding, Tian-Shi Xu et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.