Thermal-Aware GPU Energy Optimization for On-Device LLM Split Fine-Tuning
Split fine-tuning divides a large language model (LLM) between a device and an edge server at the base station, with the front layers on the device and the remaining layers on the server. However, the device's limited thermal dissipation can easily cause graphics processing unit (GPU) overheating, which in turn increas...