When training large language models with reinforcement learning, terminal rewards provide little guidance about which steps matter. Common methods for assigning step credit overlook that work built on uncorrected mistakes is wasted while independent work remains valid. With only a final success/failure reward, every st...
Zi-Yi Chen, Yan Zhang, Jian-Hui Wei et al.· 0 citations
Recurrent Longitudinal Memory (ReLMem), a framework that learns to maintain fixed-capacity patient memory for efficient downstream prediction with a frozen LLM, is introduced and a multi-granularity optimization strategy to preserve task-relevant information throughout recurrent updates and support downstream predictio...
Zi-Jie Meng, Xi-Wei Dai, Ying-Ying Zhang et al.· 0 citations
This work develops MedUAG, an end-to-end trained unified medical model that achieves strong performance across a wide array of understanding and generation tasks, establishing a competitive baseline and paving the way for next-generation medical multimodal systems.
Zi-Jie Meng, Yun-Chen Zhang, Hualiang Wang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.