This analysis covers 2,857 report-test-patch triplets from Defects4J and SWT-Bench using two widely adopted instruction-tuned LLMs from distinct model families and finds that both models exhibit systematic optimism relative to humans and only modest rank agreement, motivating bias-aware evaluation.
Wendkûuni C. Ouédraogo, Yinghua Li, Xueqi Dang et al.· arXiv.org· 0 citations
This work introduces L\"etzCross, a benchmark for cross-lingual page-level retrieval over Luxembourgish PDF documents, with document pages indexed as images and queries provided in English, French, German, and Luxembourgish, and examines single-language and multilingual fine-tuning.
Omar El Bachyr, Fred Philippy, Laura Bernardy et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.