A Dynamic Evaluation Framework for LLM Instruction Following: Multi-dimensional Verification and Iterative Feedback
Abstract
Accurately evaluating the instruction-following ability of Large Language Models (LLMs) is crucial for their practical deployment. Existing evaluation methods mainly rely on static assessment of single-pass generations, making it difficult to comprehensively measure instruction-following performance or models’ capability for self-correction under complex instructions. Moreover, current benchmarks focus primarily on formal constraints while overlooking intention understanding and semantic quality. To address these limitations, we propose Dynamic Instruction-Following Evaluation Framework (DIFE) which is a dynamic evaluation framework that integrates multi-dimensional verification with iterative feedback refinement. Specifically, we construct a three-dimensional evaluation taxonomy covering intention understanding, formal compliance, and semantic quality, together with a hybrid verification mechanism combining rule-based validation and LLM-as-a-Judge assessment. Based on the verification results, the framework automatically generates fine-grained feedback to guide iterative response refinement and reevaluation. Experiments on multiple mainstream LLMs show that the proposed framework provides more comprehensive diagnosis of instruction-following failures and consistently improves instruction-following success rates through iterative refinement, offering an effective closed-loop solution for both evaluating and enhancing LLM instruction-following capability.