TokenCast: Forecasting Token Consumption During LLM Agent Execution
TokenCast is proposed, which learns a composable cost representation for each execution segment, recording its own consumption and the context growth it introduces, and captures the extra input cost incurred when context from earlier segments is re-read by every later call.
Chao-Qian Ouyang, Ling Yue, Li-Bin Zheng et al.
· 0 citations