Mobility-Aware Vehicle Selection and Bandwidth Allocation for Federated Learning in Vehicular Networks: An HPPO Approach
Federated learning (FL) has emerged as a promising paradigm for enabling distributed model training in vehicular networks while keeping raw data local. However, the dynamic mobility of vehicles and the limited spectrum resources create critical challenges for efficient FL execution. In particular, due to the high mobility of vehicles, the candidate set of participating vehicles keeps changing, which makes fixed-threshold selection strategies difficult to be applied effectively. Moreover, vehicle mobility also causes time-varying channel status, and if resource allocation is performed only once at the beginning of each training round, it might not match the varying channels, resulting in imprecise resource allocation and accordingly low resource utilization. To address these issues, we propose a dynamic vehicular FL framework where the long time duration is discretized into fine-grained time slots. A long time sequence two-timescale optimization problem is then formulated to jointly conduct vehicle selection and slot-level communication bandwidth allocation. To solve it, we design a hierarchical Markov Decision Process (H-MDP) framework, and then develop a hierarchical Proximal Policy Optimization-based vehicle selection and bandwidth allocation (HPPO-VSBAFL) strategy, consisting of two cooperative agents: a vehicle selection agent (VSA) for round-level participant selection, and a bandwidth allocation agent (BAA) for slot-level spectrum allocation. Extensive experimental results based on CIFAR-10 with ResNet-18 demonstrate that the proposed HPPO-VSBAFL framework significantly improves FL accuracy compared to baseline schemes, and can effectively adapt to the highly dynamic vehicular environments.