We present AgentLens, a production-assessed benchmark for interactive code agents. Most code-agent benchmarks reduce a run to a single bit -- did the task pass? -- but the people who actually use these agents experience the entire trajectory: how the agent follows instructions, uses its tools, verifies its own work, recovers from mistakes, and talks to them along the way. AgentLens evaluates that whole trajectory. It pairs formal verification, where an objective check exists, with LLM-written trajectory reviews and side-by-side comparisons, so that each run yields a readable explanation of why the score is what it is. This makes AgentLens useful for more than ranking models: we use it to diagnose model behavior, compare successive versions of our own agent, and catch product regressions in a nightly evaluation pipeline. We release the benchmark as open source at https://github.com/agent-lens/agent-lens-bench.
Andrey Podivilov, Vadim Lomshakov, S. Savin et al.· 2 citations
It is argued that scientific memory should be evaluated as budgeted, modality-aware context restoration rather than as an unconstrained architecture leaderboard, and the datasets, harness, raw outputs, judgments, and scripts are released to reproduce the results and serve as tools for such evaluation.
Maksim Sheverev, David B. Finkelstein, Sergey I. Nikolenko· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.