Jul 2026
Flat Score, Amplified Failures: How the Error Budget Masks Damage in Quantized LLM Agents
This claim for multi-turn, tool-calling agents, where it now matters most, is tested for post-training quantization to 4-bit weights and diagnostics, the per-channel error rate and success under a shrinking budget come from logs benchmarks already collect.
Jiwon Jang, Kisu Yang, Heuiseok Lim et al.
· arXiv.org · 1 citation