Long-context LLMs can now ingest entire online discussion threads, but understanding their social discourse requires more than reading a long document: models must track parent-reply relations, turning points, scoped subtrees, cross-branch contrasts, and participant trajectories. To test this structure-aware social rea...
Xin-Yi Liu, R. Khaziev, Dilek Hakkani-Tur et al.· 0 citations
The PatientAgentBench framework is released as a reproducible, clinician-validated evaluation standard to help the field close this gap in healthcare agentic healthcare, and is validated by licensed clinicians annotated shared conversations.
Korosh Vatanparvar, Ashutosh Joshi, Maria Xenochristou et al.· arXiv.org· 0 citations
IMCBench is introduced, an image-grounded, multi-turn medical conversation benchmark that pairs real, publicly available clinical images with synthetic patient profiles to simulate realistic patient-clinician interactions and demonstrates that accurate clinical description does not guarantee safe patient guidance, moti...
Maria Xenochristou, Ashutosh Joshi, Korosh Vatanparvar et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.