Thermal-Aware GPU Energy Optimization for On-Device LLM Split Fine-Tuning
Abstract
Split fine-tuning divides a large language model (LLM) between a device and an edge server at the base station, with the front layers on the device and the remaining layers on the server. However, the device's limited thermal dissipation can easily cause graphics processing unit (GPU) overheating, which in turn increases energy consumption. In this paper, we propose a thErmal-aware GPU eneRgy Optimization scheme, termed EURO, to reduce the device's energy consumption during LLM split finetuning. Specifically, the EURO employs an adaptive batch size scheduling policy to prevent the device's GPU from overheating, which dynamically reduces the batch size as the GPU temperature approaches the thermal threshold. In addition, the EURO uses a GPU frequency scheduling strategy to determine the optimal GPU frequency that minimizes on-device energy consumption for each batch size. Extensive experiments on our built testbed demonstrate that the proposed scheme can reduce the device's energy consumption by at least $\mathbf{1 5. 2 6 \%}$ compared with state-ofthe-art baselines while effectively avoiding GPU overheating.