Preprint
Aug 2026
HarnessOpt-Bench: Evaluating LLMs at Harness Optimization
This work evaluates 5 frontier LLMs as optimizers both under a shared coding harness and under their native harnesses across 4 downstream tasks, and establishes harness optimization as a measurable and discriminative capability with large space for improvement.
Varun Ursekar, Apaar Shanker, Yash Maurya et al.
· 1 citation