AgentPerfBench: A Benchmarking and Evaluation Suite for Inference Performance of Agentic LLMs
The chat-to-agentic gap in current inference benchmarks is quantified and per-kernel GPU resource utilisation via roofline analysis is characterised via roofline analysis.