This work proposes soft-target fine-tuning (SoFT) to balance learning from teacher demonstrations with retaining the Base model's existing capabilities, with improvements in both in-distribution capability acquisition and out-of-distribution generalization.
Hui-Hao Jing, Wen-Bin Hu, Shao-Jin Chen et al.· 0 citations
This survey treats isolation as a first-class principle for LLM-agent system safety, and organizes the literature with a boundary-centric taxonomy of five boundaries: user-agent, agent-tool, agent-execution, agent-agent, and system-environment.
This work proposes RLPF, reinforcement learning from performance feedback, which turns execution outcomes into a staged reward, and suggests that code agents can be trained not only to pass tests, but also to optimize the programs they write.
Huihao Jing, Hao-Zhe Cui, Wenbin Hu et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.