LoRA fine-tuning adapts small language models (SLMs) to heterogeneous instruction data within a low-rank update subspace, making it vulnerable to three structural problems: conflicting gradients that cancel, static data selection that cannot track evolving learning dynamics, and subspace saturation that causes later up...
Hong-Yu Cao, Yan-Chi Liu, Kun-Peng Liu et al.· 0 citations
PolicyMem is introduced, a geometric policy memory that externalizes natural-language policies as reusable geometric memory objects represented by low-rank subspaces in a shared representation space that achieves state-of-the-art unsafe behavior detection while enabling effective policy attribution, rewriting, and post...
Yuan-Chen Bei, Zheng-Zhang Chen, Yan-Jun Zhao et al.· 0 citations
This work investigates the problem of LLM knowledge updates, which requires simultaneously unlearning unwanted information and learning new knowledge, and proposes LOKA, a conflict-aware framework for Large language mOdel Knowledge updAtes.
Binchi Zhang, Zhengzhang Chen, Zaiyi Zheng et al.· Annual Meeting of the Associ...· 0 citations
UMCTS is an Uncertainty-aware Monte Carlo Tree Search framework that combines the language understanding capability of large language models with the reliability of well-established solvers and achieves state-of-the-art solution accuracy and improves efficiency by reducing token usage.
Linlin Yu, Xujiang Zhao, Dong Li et al.· Annual Meeting of the Associ...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.