Review
Aug 2026
SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring
SWE-Bench ProMax is introduced, an expert-curated, multilingual code refactoring benchmark of 170 instances drawn from real commits across seven programming languages, which presents a meaningful and unsaturated challenge for current AI coding agents.
Yuling Shi, Jingheng Xu, Kelin Fu et al.
· 5 citations