Characterization of Request and Token Energy Costs for LLM Inference Workloads on GPU Platforms
Results show that energy-aware serving should jointly optimize both request energy and token energy, rather than only reducing per-token energy cost, and substantially narrows the dense-vs-MoE token-energy gap.
P. Vellaisamy, Vanessa Lam, Shawn Blanton et al.
· 0 citations