Learning-Guided Task Refinement for Multi-UAV Swarm Coordination
Abstract
Multi-UAV swarm coordination requires high-level task refinement across heterogeneous game-like operation segments with different controllers, risks, and time-dependent rewards. Existing hand-written refinement rules are difficult to tune when an intermediate action has no direct reward but changes downstream losses and completion time. This paper presents a simulation-grounded learning-guided refinement framework for swarm coordination. RAE-style symbolic refinement first enumerates symbolically feasible object options, the swarm simulator evaluates their executed operation utility, and a contextual Learn-π policy distills these evaluations into deployable method and order preferences. In a controlled EMI-density sweep, Learn-π achieves zero mean held-out regret over 30 cases by selecting strike-first or suppress-first according to continuous jam coverage, whereas fixed-rule and context-free selectors cannot. On 50 randomized held-out maps it further attains the lowest zero-rollout regret among deploy-time selectors and degrades gracefully under Gaussian sensor noise on the jam feature. Full-operation and multi-scale studies show that the learned policy completes heterogeneous tasks in the 16-UAV setting and adapts visit order when airframes are scarce.