Skip to content
Conference

Deep Reinforcement Learning-based Dynamic TWT Scheduling for Heterogeneous Wi-Fi Networks

Jul 2026 · International Conference on Computer Communications and Networks · pp. 1-9 · 0 citations · 22 references

Abstract

Target Wake Time (TWT) in IEEE 802.11ax enables significant power savings by coordinating station sleep schedules, but optimal TWT parameter selection under heterogeneous traffic remains an open challenge. We present a deep reinforcement learning framework for dynamic TWT scheduling that is protocol-compliant by design. The agent’s observation is restricted to 802.11ax buffer status reports, 802.11k station statistics, and AP-derived TWT metrics, ensuring direct deployment without firmware modifications. The system employs Proximal Policy Optimization (PPO) over a factored MultiDiscrete action space of 380 schedule-assignment combinations, trained end-to-end within ns-3 via shared-memory with episode-level process isolation for stable, reproducible training. We evaluate two architectures, MLP-PPO and LSTM-PPO, against an analytical M/D/1 baseline across three reward presets (energy, throughput, queue) over 50 seed-disjoint episodes with 16 stations from four heterogeneous device classes. Both RL agents substantially outperform the analytical baseline on all presets. On the energy preset, MLP-PPO reduces aggregate energy by 44% while improving throughput by 18% and reducing drops by 73%. On the queue preset, LSTM-PPO achieves 15% higher reward with 77% fewer drops by leveraging temporal correlations across beacon intervals. The two architectures exhibit complementary strengths. MLP-PPO excels on the energy-dominated preset where a stationary policy suffices while LSTM-PPO’s recurrent state captures multi-step queue-drain dynamics that transfer across network configurations, motivating future ensemble approaches.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.