PhysAssistBench is introduced, a benchmark for interactive doctor-patient-EHR assistance that uses a scalable pipeline to construct agentic patients: interactive, record-grounded agents that turn static EHR records into multi-turn clinical scenarios while preserving clinical factuality.
T. Du, Peijie Yu, Sihan Shang et al.· arXiv.org· 0 citations
Findings suggest that AI can be easily persuaded by what people say, who says it, and how the opinion is presented, enabling its safe and reliable use in high stakes medical decision making.
Jiayuan Zhu, Jiazhen Pan, Feng-Lin Liu et al.· 0 citations
Under matched backbone and inference settings, J-CoT-Zero matches or exceeds the strongest evaluated latent-reasoning baseline on every benchmark, while J-CoT-Train obtains the highest score across the evaluated mathematical, scientific, coding, and structured path-reasoning tasks.
Jun-De Wu, Jiayuan Zhu, Feng-Lin Liu et al.· arXiv.org· 1 citation
A Dynamic, Automatic and Systematic red-teaming audit framework that continuously stress-tests LLMs for health across four safety-critical axes: robustness, privacy, bias and hallucination, which provides a scalable framework for surfacing latent risks before such systems are deployed in consumer-facing health assistants and broader clinical workflows.
Jiazhen Pan, Bailiang Jian, Paul Hager et al.· Nature Health· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.