Large language model (LLM) agents are vulnerable to safety risks such as injected malicious instructions or misleading information, motivating runtime defenses that prevent unsafe action in execution across diverse risks while preserving benign-task utility. Existing system-level defenses either focus on risk detection...
Zhuo Liu, Mo-Xin Li, Zhi-Xin Ma et al.· 0 citations
A novel set identifier paradigm is introduced, representing each item as a set of order-agnostic tokens, which proves SETRec’s superior efficiency and scalability on cold-start items as model sizes increase, and SETRec++’s potential in the scalability of order-agnostic identifier.
Xin-Yu Lin, Chuan-Bo Zhang, Yu-Fan Liu et al.· ACM Transactions on Recommen...· 0 citations
It is demonstrated that the hidden states of probed answers more effectively differentiate distinct solution paths than semantic embeddings, and the perplexity of probed answers serves as a practical proxy for reasoning correctness.
Yi Fang, Quek Shen, Chengping Li et al.· 0 citations
A novel framework that finetunes generative models using distribution-wise rewards, ensuring better alignment with real-world data distributions is presented, and a subset-replace strategy that efficiently provides reward signals by updating only a small subset of a generated reference set is introduced.
Ruihang Li, Mengde Xu, Shuyang Gu et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.