Semantic Drift in Bug Resolution: How Behavioral Signals Propagate from Reports to Tests and Patches
This analysis covers 2,857 report-test-patch triplets from Defects4J and SWT-Bench using two widely adopted instruction-tuned LLMs from distinct model families and finds that both models exhibit systematic optimism relative to humans and only modest rank agreement, motivating bias-aware evaluation.