Adaptive rocket trajectory optimization via reinforcement learning and FPGA acceleration
This paper presents an integrated framework for reusable rocket trajectory optimization that combines deep reinforcement learning with FPGA-based inference acceleration. The proposed Adaptive Hardware-Accelerated Twin Delayed Deep Deterministic Policy Gradient (AHA-TD3) framework couples a hierarchical control policy, an online adaptation mechanism, and an FPGA-oriented inference pipeline. Within the simulation and board-level hardware validation considered in this study, AHA-TD3 improves landing success rate, position accuracy, and fuel consumption relative to the compared PPO, DDPG, SAC, TD3, and SCP-MPC baselines. The FPGA implementation achieves up to 12.7× lower inference latency than the CPU software baseline while maintaining low power consumption. These results indicate the potential of combining adaptive reinforcement learning and hardware-software co-design for real-time reusable launch vehicle guidance, while further high-fidelity hardware-in-the-loop and flight-oriented validation remain necessary before operational use.