Preprint
Aug 2026
Joint Optimization of Tool Creation and Use for Large Language Model Agents
A reinforcement learning framework that jointly trains tool creation and tool use inside a single policy, with three separate reward axes that catch schema, code, and outcome failures independently, so each failure mode contributes its own gradient.
Zhi Rui Tam, Chieh-Yen Lin, Yun-Nung Chen et al.
· 0 citations