Review
Jul 2026
Programmers Are Poor and Overconfident Judges of LLM-Generated Assertions
The findings suggest that, contrary to common assumptions, AI assistance may not improve the reliability of code comprehension and review, and highlight the importance of helping developers evaluate machine-generated reliability artifacts, in addition to generating them.
Zhanna Kaufman, Yuriy Brun, Adithya Murali et al.
· arXiv.org · 0 citations