LLM agents accumulate interaction histories that grow linearly with task length, causing quadratic inference cost scaling and performance degradation from attention dilution. Existing context-compression methods learn what to discard offline: by contrastively optimizing guidelines, distilling compressors, or training c...
Shantanu Dixit, Anson Bastos, Xu-Chao Zhang et al.· 0 citations
Pseudo Self-Distillation is presented, a framework that enables small language models to construct hierarchical memory representations by distilling behavior from a strong black-box oracle through a multi-stage training pipeline, with off-policy PSD achieving the strongest results across most conditions.
Pirzada Suhail, Meng-Lin Xia, Xu-Chao Zhang et al.· 0 citations
This work introduces WebXSkill, a framework that bridges a grounding gap with executable skills, each pairing a parameterized action program with step-level natural-language guidance, and finds that better skill deployment mode depends on a model's plan and execution capability.