Skip to content

AgentPerfBench: A Benchmarking and Evaluation Suite for Inference Performance of Agentic LLMs

Sep 2026 · 0 citations · 16 references
Computer Science

TL;DR

The chat-to-agentic gap in current inference benchmarks is quantified and per-kernel GPU resource utilisation via roofline analysis is characterised via roofline analysis.

Abstract

The optimization of LLM serving engines, such as vLLM and SGLang, is largely benchmark-driven: optimizations, scheduling policies, hardware and system designs are all selected based on representative workloads. However, a significant mismatch has emerged in the agentic era. Existing benchmarks primarily focus on simple single-turn chatbot workloads. LLM applications are increasingly agentic: coding agents, terminal execution systems, and tool-use agents issue multi-turn requests with growing context lengths. We introduce AgentPerfBench, a benchmark suite for agentic inference. It uses real traces from agentic benchmarks, such as SWE-Bench and TerminalBench, alongside standard chat baselines. This enables benchmarking of models on multi-turn tasks involving tool calling, skill utilization, and increasing context lengths. AgentPerfBench also samples from empirical distributions of input length, output length, and turn count derived from the real traces, generating representative synthetic profiles for cheap and accurate measurements on new hardware. In addition, we further find that several existing benchmarks fail to accurately reflect real hardware performance for two key reasons: 1) they do not account for realistic context-length growth, and 2) they measure inference performance without operating at hardware saturation. We discuss these issues in detail and provide rich kernel-level Nsight Compute (NCU) traces to construct a new multi-dimensional roofline model that captures hardware-system limitations in both memory bandwidth and memory capacity footprint. The benchmarking suite then includes automated scripts to identify potential bottleneck conditions on emerging hardware when evaluated with diverse agentic traces. Together, these contributions quantify the chat-to-agentic gap in current inference benchmarks and characterise per-kernel GPU resource utilisation via roofline analysis.

View source

Similar papers

Preprint Aug 2026

From LLM Inference to Agentic Workloads: Characterization and Implications for Serving Systems

This work presents AgentSysBench, a benchmark suite and measurement toolkit with ten representative agentic applications and unified systems-level instrumentation, and identifies six properties that distinguish agentic workloads from conventional LLM serving.

Chaokun Chang, Yu-Kun Zhou, Kai-Hua Fu et al. · 11 citations
Preprint Aug 2026

LLM4LLM: Bridging Kernel Benchmarks and Real Deployment via Closed-Loop Agentic Optimization

LLM4LLM is introduced, a deployment-aware closed-loop optimization framework that starts from a target inference script, extracts phase-aware optimization tasks, searches with an experience-guided episodic agent, and accepts patches through in-model validation.

Hui Zeng, Pengfei Yang, Yanxin Chen et al. · 0 citations
#artificial intelligence Preprint Sep 2026

TicTacBench: Benchmarking Timing Closure Capabilities of Coding Agents

Recent advances in large language models (LLMs) have led to the emergence of coding agents capable of performing complex engineering tasks, including register-transfer level (RTL) design and optimization. Existing RTL benchmarks mainly evaluate functional correctness and performance, power, and area (PPA) of the genera...

Bo-Wei Wang, Zhigang Fang, Zhijie Yang et al. · 0 citations
Preprint Oct 2026

Engineering Sustainable Agents: A Systematic Comparison of Agentic LLMs for Developer Workflows

Large language models (LLMs) are increasingly used in software engineering, including agentic systems that coordinate multiple agents, but impose higher computational and environmental costs. In this paper, we present a comprehensive empirical study of agentic LLM systems across five software engineering tasks: code ge...

Merve Astekin, Y. N. Tun, Arda Goknil et al. · 0 citations
#artificial intelligence Preprint Sep 2026

DAREBench: Deployment-Aware and Reliable Evaluation of Models as Agents

DAREBench (Deployment-Aware and Reliable Evaluation of Models as Agents), a benchmark designed to capture workload variation and support reliable agent evaluation, is introduced, suggesting that agent deployment and model selection should consider workload profiles, deployment mode, and accuracy--cost trade-offs rather...

Yu Liu, Zhi-Lin Liu, Zhi-Wei Yang et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Oct 7, 2026

Discovering the value of humanistic inquiry

Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.

Microsoft Research Blog Oct 7, 2026

Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses

Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.