An embodied kitchen assistant must do more than recognize food in isolated frames. It must track ingredient states over time and integrate visual observations with recipe and nutritional knowledge to support constraint-aware decision-making. We formalize this capability as \emph{Embodied Nutrition Management}: perceivi...
Yue-Ling Wei, Xiang-Chen Wang, Jian-Hui Pan et al.· 0 citations
JustMem is introduced, which stores conversation history as compact atomic memories and adapts memory access along two dimensions to each query and achieves the highest mean accuracy and retrieval recall among the compared memory systems while using substantially fewer generative-model tokens for memory construction an...
Guan-Hua Chen, Yan-Ting Wang, Wen-Jing Zhi et al.· 1 citation
SAFT (Safety-preserving Adaptation via Fine-tuning Transfer), a safety-preserving adaptation framework that decouples task learning from alignment preservation by learning a safety-guided task update on the paired pretrained base model, rectifying task gradients to avoid conflicting directions with respect to a safety...
Zhiwen Ruan, Yan Yang, Zhuocheng Liang et al.· Proceedings of the 32nd ACM...· 0 citations
Long-term language-model agents rely on external memory across interactions. Atomic memories are particularly useful: their fine-grained semantic boundaries enable precise retrieval and direct comparison between observations. Yet accumulating atoms inevitably become redundant, overlapping, or conflicting. Existing meth...
Jianjie Zheng, Peng Lai, Sijie Cheng et al.· 0 citations
Reinforcement learning (RL) excels on tasks with verifiable rewards, but in open-ended tasks, the reliability of reward models remains a key challenge. Existing solutions either depend on costly proprietary LLM-as-a-Judge systems or opaque scalar reward models that lack interpretability. Recent works on generative rewa...
Peng Lai, Yi-Chao Du, Junchao Wu et al.· 1 citation
AlignDiff, a preference data filtering framework driven by intrinsic model signals, first identifies samples with clear preferences using both positive and inverse signals, then prioritizes the more challenging samples based on the average negative log-likelihood gap, encouraging the model to learn richer information f...
This survey provides a comprehensive analysis of reasoning economy in both the post-training and test-time inference stages of LLMs, encompassing the cause of reasoning inefficiency, behavior analysis of different reasoning patterns, and potential solutions to achieve reasoning economy.
CAST (Concept-guided Artifact Suppression Tuning), an SAE-based framework for auditable clinical text classification, improves over its corresponding fine-tuned encoder baselines and remains competitive with strong LLM baselines, while producing a feature-level audit trail of the clinical concepts that support each pre...
P-Bench is built, a benchmark comprising 425 open-ended, realistic hypothesis-testing tasks spanning economics, biology, and medicine and introduces Fisher-R1, an open-weight LLM agent trained for rigorous hypothesis testing using synthetic tasks and reinforcement learning.
Jia-Cheng Miao, Jin Mu, Guan-Hua Chen et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.