Rethinking Code Performance Benchmarks for LLMs
An LLM-based multi-agent framework to generate performance-oriented tests that expose runtime differences more effectively than the original tests is proposed, which uses three separate agents to generate, diagnose, and repair deterministic tests that preserve functional correctness while better exposing performance differences.