This work proposes soft-target fine-tuning (SoFT) to balance learning from teacher demonstrations with retaining the Base model's existing capabilities, with improvements in both in-distribution capability acquisition and out-of-distribution generalization.
Hui-Hao Jing, Wen-Bin Hu, Shao-Jin Chen et al.· 0 citations
Evaluated across three main benchmarks, two domain-specific studies, and six LLMs, SkillRevise substantially outperforms one-shot baselines, and the revised skills transfer across both executors and task environments, suggesting that SkillRevise captures reusable procedural knowledge beyond any single executor.
Yuxuan Liu, Zhao-Chen Su, Lin Xie et al.· arXiv.org· 17 citations
Overall, persistent skill self-evolution is better understood as sparse, validation-filtered search with model- and benchmark-dependent returns, rather than steady improvement from additional rounds.
Yuxuan Liu, Zhaochen Su, Yuhao Zhang et al.· 2 citations
This survey treats isolation as a first-class principle for LLM-agent system safety, and organizes the literature with a boundary-centric taxonomy of five boundaries: user-agent, agent-tool, agent-execution, agent-agent, and system-environment.