Intent drift is established as a measurable multi-turn failure mode and explicit state maintenance as a partial mitigation and IntentFlux is introduced, an executable benchmark that converts verifiable tasks into dialogues with controlled intent changes while preserving their original graders.
Yan-Jie Zhang, Bo-Wen Cao, Zi-Xin Chen et al.· 0 citations
As LLMs become increasingly capable of completing tasks for users, a central concern is that everyday AI use may become primarily cognitive offloading, eroding the opportunities through which people develop their own capabilities. We analyse large-scale human-LLM conversations to ask whether informal learning behaviors...
This work presents TACT (Taxonomy-Aligned Conversational Tutor), a human-grounded framework for post-training and evaluating pedagogically adaptive ESL tutors and develops two complementary taxonomies: the Tutor-Strategy Taxonomy with 13 tutor response strategies and the Student-Move Taxonomy characterizing learner beh...
Dongjie Yang, Siya Lin, Leixian Shen et al.· 0 citations
WILDTRACE is introduced, a benchmark of 481 tasks over 214 naturally occurring long-form sources such as technical incident reports and lesser-known literary narratives, where all evidence trails arise from the document's own causal, temporal, and narrative logic.
Zixin Chen, Peng Liu, Haobo Li et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.