This work introduces PonyEval, a SWE-bench-style benchmark of 291 real GitHub issue-pull-request pairs from 15 Pony repositories that defines a matched evaluation with mini-SWE-agent 2.4.6 for GPT-5.6-sol, DeepSeek-V4-Pro, GLM-5.2, MiniMax-M3, and Kimi-K3, followed by strict patch application, compilation, and hidden-t...
Bang Xie, Hao Liu, Zhen-Yu Shi et al.· 0 citations
AppEval is presented, a benchmark and native-toolchain evaluation framework for mobile application repair across HarmonyOS/ArkTS, iOS/Swift, and Android/Kotlin, and shows that mobile repair performance depends strongly on the evaluated agent while demonstrating why runtime-aware acceptance is necessary for meaningful c...
Bang Xie, Hao Liu, Zhen-Yu Shi et al.· 0 citations
OdinEval is presented, a reproducible benchmark built from documented defects in public Odin repositories built from documented defects in public Odin repositories, that evaluates six language models on 168 filtered instances under one shared protocol.
Bang Xie, Hao Liu, Zhi-Yuan Peng et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.