A Dynamic Evaluation Framework for LLM Instruction Following: Multi-dimensional Verification and Iterative Feedback
Accurately evaluating the instruction-following ability of Large Language Models (LLMs) is crucial for their practical deployment. Existing evaluation methods mainly rely on static assessment of single-pass generations, making it difficult to comprehensively measure instruction-following performance or models’ capabili...