LoopArena: Benchmarking Models as Runtime Controllers for Loop Engineering
This work introduces LoopArena, a benchmark for evaluating how well one model can guide a separate coding agent through a long-running task, and evaluates this ability in three complementary settings that differ in execution scope and cost.
Yi Wang, Haopeng Zhang, Chen Huang et al.
· 0 citations