Preprint
Aug 2026
Understanding the Energy Scaling of Large Language Model Inference Across Context Lengths and Attention Architectures
Results show that attention mechanism is the primary factor governing how decode energy scales with context length, and model size primarily determines absolute energy consumption, while batching reduces both energy per generated token and request latency by up to 87%.
Molka Chkir, Syed Muhammad Danish, Jos Höll et al.
· 0 citations