Skip to content

FedDPRL: Communication-Efficient Differentially Private Federated Reinforcement Learning With Adaptive Gradient Compression

2026 · IEEE Transactions on Communications · Vol 74, pp. 13899-13912 · 0 citations · 55 references

Abstract

Cross-silo federated reinforcement learning (FRL) trains a shared policy across silos that must simultaneously respect a per-round uplink budget and on-device privacy. Although differential privacy (DP) and gradient compression each have mature solutions in supervised federated learning, naively stacking them in the policy-optimisation regime fails, for two structural reasons. First, whole-client DP injects noise whose ratio to the signal scales as <inline-formula> <tex-math notation="LaTeX">$\sqrt {d}/N$ </tex-math></inline-formula>, which for neural policies exceeds one at every practical <inline-formula> <tex-math notation="LaTeX">$\varepsilon $ </tex-math></inline-formula>, so the policy collapses to random. Second, a single shared compression mask cannot represent each client’s descent direction under non-i.i.d. environments, driving the compressor’s contraction constant toward zero and destroying learning. In this paper, we propose FedDPRL, a federated policy-optimisation algorithm built from two coupled designs. Specifically, a record-level (per-trajectory) Gaussian mechanism clips and noises each trajectory, lifting the effective sample size to <inline-formula> <tex-math notation="LaTeX">$R=NB$ </tex-math></inline-formula> and removing the <inline-formula> <tex-math notation="LaTeX">$\sqrt {d}/N$ </tex-math></inline-formula> barrier so that DP becomes viable on neural policies. Additionally, an error-feedback Top-<inline-formula> <tex-math notation="LaTeX">$k$ </tex-math></inline-formula> compressor with per-client masks preserves each client’s own descent direction while cutting the uplink by an order of magnitude. Experimental results across classic control, continuous control (MuJoCo), and vision with <inline-formula> <tex-math notation="LaTeX">$N=20$ </tex-math></inline-formula> clients show that FedDPRL matches dense DP-FedAvg-RL utility at 13–<inline-formula> <tex-math notation="LaTeX">$14\times $ </tex-math></inline-formula> less uplink, while a shared mask collapses to near-random.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.