Jul 2026
PERFOPT-Bench: Evaluating Coding Agents on Software Performance Optimization
The results show that optimization performance is workload-dependent rather than determined by model identity alone: no single stack dominates, and changing the agent framework can materially change the same LLM's per-task speedup profile.
YI-YING Cui, Yi Xie, Piaohong Wang et al.
· arXiv.org · 0 citations