Reusable Latent Correction (RLC) is proposed, which converts one-off natural-language guidance from a black-box LLM into persistent corrective experiences in the hidden space of an SLM, enabling the SLM to reuse LLM-derived corrections during inference without any online LLM calls.
Bo-Han Zhang, Li-Nan Yue, Weibo Gao et al.· 0 citations
Learner simulation aims to reproduce how a particular learner behaves on new tasks. Although Large Language Models (LLMs) can generate increasingly fine-grained learning behaviors, existing approaches often need to repeatedly process a growing interaction history to reconstruct the learner. This introduces additional c...
Zi-Jian Chen, Zheng Zhang, Miao Jia et al.· 0 citations
This work proposes Step-wise Training for In-context Reasoning (STIR), a model to dynamically decide when to retrieve a single logically consistent next step, just using the current problem and its intermediate state as the query.
Cheng Yang, Zhenya Huang, Liyang He et al.· Proceedings of the 32nd ACM...· 0 citations
Reinforcement learning (RL) excels on tasks with verifiable rewards, but in open-ended tasks, the reliability of reward models remains a key challenge. Existing solutions either depend on costly proprietary LLM-as-a-Judge systems or opaque scalar reward models that lack interpretability. Recent works on generative rewa...
Peng Lai, Yi-Chao Du, Junchao Wu et al.· 1 citation
SafeIMG is introduced, a safety-oriented benchmark spanning 12 public- and individual-safety scenarios generated using GPT Image 2.0 that provides human annotations that localise suspicious regions and explain local artefacts and higher-level commonsense or physical inconsistencies.
Yizhi Wang, Yichen Xiao, Linan Yue et al.· 0 citations
This work proposes Step-wise Training for In-context Reasoning (STIR), a model to dynamically decide when to retrieve a single logically consistent next step, just using the current problem and its intermediate state as the query.
Cheng Yang, Zhenya Huang, Liyang He et al.· Proceedings of the 32nd ACM...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.