Large language models deployed as personalized assistants must reason over long, evolving interaction histories. However, in long-term dialogue reasoning, relevant evidence is scattered across sessions, preferences may be revised over time, and standard long-context training fails to address these challenges under data...
Na-En Xu, Wan-Qing Cui, Yi-Bo Hu et al.· 0 citations
Safety-Flag is introduced, which places seven widely used safety benchmarks (BeaverTails, XSTest, Ethics, WildGuard, Aegis, ToxiChat, and ToxiGen) into a single balanced flag / do-not-flag protocol.
A conformal certificate can be valid when an LLM answers alone and invalid when the same LLM sees peers that unanimously assert a wrong answer. The question is unchanged; the model's score for the correct answer changes. We call this a score-mechanism shift: clean calibration certifies how the model scores answers alon...
Yi-Bo Hu, Han-Yuan Su· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.