This work proposes a two-component framework in which a profile generator summarizes a student's history and a simulator predicts student turns conditioned on the resulting profile, which trains both components with reinforcement learning (RL), yielding profiles optimized for faithful student simulation.
Zhangqi Duan, Shuyan Huang, Alexander Scarlatos et al.· arXiv.org· 0 citations
GRASP is introduced, a reinforcement learning (RL) framework for training agents to adaptively coordinate complementary retrieval tools during multi-step reasoning, and it is suggested that learning to coordinate retrieval signals and context granularity is critical for agent's correct reasoning.
An automated, data-driven approach to uncover patterns, which the authors term traits, of effective human-AI interaction that are aligned with task outcomes is explored and Principal Trait Analysis is proposed, a Principal Component Analysis-inspired algorithm for deriving common traits from patterns in LLM conversations.
Hunter McNichols, Kai Du, Andrew S. Lan· 0 citations
This work proposes a framework for extracting reusable developer preferences from interaction traces, generates personalized skills through rule-based bootstrapping and evidence-grounded refinement, and evaluates them using a reproducible replay framework with an interactive, trajectory-conditioned LLM-based human developer simulator.
Shuyan Huang, Kai Du, Andrew S. Lan· 3 citations· ⚡2
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.