tokenwatt: ground-truth energy measurement for LLM inference across heterogeneous CPU power islands
A RAPL-based measurement harness for large language model inference on heterogeneous CPUs. Separates prefill from decode, records KV cache growth at every decode step, subtracts a verified idle baseline, and degrades to timing-only measurement rather than estimating energy when the counters are unreadable. Includes the...