This work proposes a novel training framework that explicitly aligns the LLM-generated judgment distribution with human evaluation distributions, and incorporates adversarial training to ensure a robust alignment with this true distribution, rather than overfitting to its imperfect approximation.
Lu-Yu Chen, Zeyu Zhang, Hao-Ran Tan et al.· Neural Information Processin...· 4 citations· ⚡1
It is argued that, in dynamic long-horizon interactions, memory is not a static collection of facts but a lifecycle of explicit operations, including remembering, forgetting, updating, reflecting, and their compositions, which reveal that current systems remain far from uniformly reliable.
Xixuan Hao, Zeyu Zhang, Zehao Lin et al.· arXiv.org· 3 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.