How Powerful are LLMs in Generating Formal Program Specifications?
Coins is introduced, a Rocq based evaluation framework that assesses specification quality by instantiating specifications under evaluation on trusted test cases and generating concrete proof obligations, and finds that accurate specification evaluation, rather than model scaling alone, is central to understanding the power of LLMs for specification synthesis.