Modern LLM post-training composes supervised fine-tuning (SFT), reinforcement learning with verifiable rewards (RLVR), and on-policy distillation (OPD) into multi-stage pipelines, yet these stages are typically designed and evaluated in isolation. We show that this composition is consequential: a stage that improves th...
Emre Can Acikgoz, Yang Li, Z. Liu et al.· 0 citations
This work introduces AGNI, an automated pipeline that extracts trajectory-relevant assumptions, injects targeted environmental changes, and validates that the resulting novel tasks remain solvable and highlights a gap between task competence and adaptive capability and motivate environmental variation as a core dimensi...
Janvijay Singh, Vaishnavi Shrivastava, Dilek Hakkani-Tur et al.· 0 citations
ReasoningFlow is introduced, a framework that captures the discourse structures of LRM reasoning traces into fine-grained directed acyclic graphs (DAGs) and reveals diverse fine-grained reasoning behaviors that can be used for better reasoning trace monitorability.
Hear2Act is introduced, a unified evaluation protocol for text and spoken assistants with 480 persona-grounded scenarios, hidden user concerns, and objectively verifiable outcomes that show that prosody matters when lexical evidence is insufficient, and that audio-capable LLMs can recover information from speech but do...
Xin-Yi Liu, H. Nayyeri, Dilek Hakkani-Tur et al.· 3 citations· ⚡1
Gradient-Based Connections (GBC) is proposed, an approach for fine-grained attribution and optimization of multi-agent systems that improves multi-agent performance and outperforms strong single-agent and multi-agent baselines and higher attribution quality is associated with greater optimization effectiveness.
Xiaocheng Yang, A. Alrabah, Dilek Hakkani-Tur et al.· SIGDIAL Conferences· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.