StructureBench: A Unified Benchmark Suite for Multi-Scenario Structured Generation Tasks with On-Device Models
The experiments show that constrained decoding consistently enforces syntactic validity, but does not reliably improve semantic accuracy and may even degrade performance for smaller models or complex grammars, and reveal clear task-and model-dependent boundaries for effective constrained decoding.