SLMs as Multi-Agent Routers: A Progressive SFT and Reinforcement Learning Approach
A small language model is trained via supervised fine-tuning followed by reinforcement learning to jointly perform agent selection and structured parameter generation for downstream tool calls, using a hierarchical reward function grounded in retrieval relevance along with query-agent topic alignment to learn task-dependent agent suitability from retrieval performance.