Skip to content

Author

A. Tsourdos

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access Jul 2026

Communication-Aware Decentralised Multi-Agent Reinforcement Learning Framework for UAV-Based Wildfire Suppression: Challenges Under Realistic Communication Constraints

The increasing frequency and intensity of wildfires has created an urgent demand for scalable and autonomous wildfire response systems. While recent advances in multi-agent reinforcement learning (MARL) have demonstrated promise for collaborative uncrewed aerial vehicle (UAV)-based wildfire suppression, most existing approaches rely on simplified fire propagation dynamics and highly centralised learning architectures that are difficult to deploy in realistic operational settings. This paper presents a decentralised MARL framework for wildfire suppression that combines stochastic wildfire propagation, wind-driven spread dynamics, and communication-aware multi-agent coordination. The proposed framework extends an existing probabilistic wildfire environment through the incorporation of wind speed and directional effects, producing highly asymmetric and stochastic wildfire behaviour that more closely resembles real wildfire propagation. A decentralised Deep Q-Network (DQN) architecture is then introduced in which UAV agents learn independently through individual replay buffers. To mitigate the sparse-learning challenges introduced by decentralisation, selective experience sharing based on the SUPER algorithm is incorporated, enabling agents to exchange only high-value experiences under realistic communication constraints. Experimental results demonstrate that selective communication significantly improves containment performance and learning efficiency while preserving decentralised execution. The work highlights both the feasibility and challenges of realistic UAV swarm coordination for wildfire suppression, particularly the trade-offs between communication bandwidth, environmental stochasticity, and collaborative performance.

S. Cartwright, Maxime Collignon, Adolfo Perrusquía et al. · 0 citations
Conference Jul 2026

Recurrent Actor-Critic Navigation of Unmanned Aerial Vehicles in Arbitrary-Shape Obstacle-Dense Environments via Temporally-Augmented Observation

Reactive navigation of Unmanned Aerial Vehicles (UAVs) in cluttered environments using only local sensor feedback is severely hampered by partial observability: irregularly-shaped obstacles can disappear from the sensor’s instantaneous field of view, causing wrap-around collisions and local-minima traps. However, existing Reinforcement Learning approaches only partially address these challenges, as they commonly assume regularly shaped, often circular or convex, obstacles. This work proposes a lightweight recurrent extension of the Twin Delayed Deep Deterministic Policy Gradient (TD3) algorithm that tackles navigation in cluttered environments populated with arbitrary-shape obstacles. By augmenting the agent’s observation with a sliding window of recent states, the resulting Temporally-Augmented Observation TD3 (TAO-TD3) feeds stacked observations through a recurrent feature backbone, enabling the policy to better infer obstacle structure from temporal information. Benchmark evaluations across five environments with increasing obstacle density and count demonstrate that TAO-TD3 consistently reduces collisions and improves goal-reaching rates relative to memory-less Reinforcement Learning.

Gabriele Gemignani, Adolfo Perrusquía, Lorenzo Pollini et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.