Agent in a Bottle: Can LLM Agents Turn Their Capabilities Into Cheap, Scalable Artifacts?
This work introduces BOTTLED, a benchmark in which agents receive an entire unlabelled workload and must complete it under fixed time, compute and LLM API budgets, and finds that strong zero-shot task performance does not reliably translate into strong bottling capabilities.
Ankit Sonthalia, Haritz Puerto, Alexander Rubinstein et al.
· 0 citations