Skip to content

StructureBench: A Unified Benchmark Suite for Multi-Scenario Structured Generation Tasks with On-Device Models

· 0 citations · 25 references

TL;DR

The experiments show that constrained decoding consistently enforces syntactic validity, but does not reliably improve semantic accuracy and may even degrade performance for smaller models or complex grammars, and reveal clear task-and model-dependent boundaries for effective constrained decoding.

View source

Similar papers

#small language model Preprint Aug 2026

SchemaGUI: A Schema-Driven Benchmark for Controllable GUI Generation Evaluation

SchemaGUI, a template-based benchmark for controllable GUI generation evaluation, synthesizing paired natural language instructions and deterministic function-call references from parameterized interface schemas can generate thousands of deterministically annotated tasks in seconds without human labeling.

Jiarui Dong, Yin Cai, Zhouhong Gu et al. · 0 citations
Jul 2026

MV-Bench: Benchmarking Multimodal Large Language Models for Coordinated Multi-View Interface Construction

It is shown that current MLLMs can reproduce visual appearance but remain limited in generating the data semantics and interactive logic required by coordinated multi-view interfaces, and Iterative refinement improves code executability but does not substantially reduce the gap in data binding and interaction generation.

Yue Zhao, Hongxu Liu, Feiyu Wang et al. · 0 citations
Preprint Aug 2026

Route-Align-Verify for Functional Correctness in Code Generation

The results indicate that functional correctness in code generation can be meaningfully improved without modifying the backbone architecture, by jointly optimizing how tasks are prompted, how the model is adapted, and how final outputs are selected.

Erxue Zhou, Jing Meng, Aofan Liu · 0 citations
Preprint Aug 2026

NoTB: Oracle-Free Triage of LLM-Generated RTL via Cross-Model Formal Consensus

NoTB is introduced, an oracle-free triage framework that infers correctness from cross-model formal consensus and demonstrates that formal cross-model agreement provides a reliable basis for high-confidence triage without model-dependent oracles.

Elisavet Lydia Alvanaki, Je Yang, Biruk B. Seyoum et al. · 0 citations
Jul 2026

IFHierBench: Hierarchical Instruction Following for Large Language Models

IFHierBench is introduced, a hierarchical instruction-following benchmark of 600 prompts stratified across four constraint-tree depths and 35 distinct constraints, each prompt paired with a deterministic checker that verifies satisfaction at every scope.

Yuetian Mao, Chunyang Chen · 0 citations
Jul 2026

VClare: Resolving Imperfect Specifications in LLM-Based Verilog Generation

The first systematic study of imperfect specifications is presented and an automated framework to repair them to enhance the quality of resulting Verilog design is proposed, demonstrating the capabilities of specification repair by {VClare} as well as further potential of LLMs in front-end hardware design.

Zhuorui Zhao, Bing Li, Yu Li et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.