Memory-augmented reinforcement learning strengthens LLM agents'ability to solve complex long-horizon tasks. Skills are one such form of memory, pairing instructions with an applicability condition over task types. However, retaining every skill indiscriminately as the policy improves lets obsolete or harmful entries ac...
Yuyao Ge, Yi-Wei Wang, Yu-Chen He et al.· 0 citations
A user-centric framework for systematically auditing system prompts in AI systems, AISPA is introduced, a user-centric framework for systematically auditing system prompts in AI systems that examines specific parts of a system prompt and evaluates them along eight dimensions that matter to users.
Xiangning Lin, Shenzhe Zhu, Shu Yang et al.· arXiv.org· 0 citations
This work introduces Audio-Zero, the first label-free self-evolution framework in the field of LALMs that improves fine-grained auditory perception and reasoning and reveals that increasingly fine-grained auditory descriptions emerge naturally from game pressure.
Siqian Tong, Xuan Li, Chao-Zhuo Li et al.· arXiv.org· 1 citation
This work introduces Polarity-Prompt Contrastive Decoding (PopCD), a test-time behavior control method that generalizes contrastive decoding to broader enhancement settings and is applicable to both LLMs and Vision-Language Models without additional training.
Bao-Long Bi, Yuyao Ge, Shenghua Liu et al.· IEEE Transactions on Pattern...· 0 citations
LP-SFT, a Local-Preserving Supervised Fine-Tuning objective designed to explicitly protect this inherent entropy structure, improves overall performance over vanilla SFT and recent SFT-enhancement baselines, suggesting that local preservation helps mitigate capability degradation without collapsing sampling-accessible...
Yueyang Wang, Baolong Bi, Shuo Lu et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.