GPAgentBench-2K: Benchmarking Large Language Model Agents in Complex Clinical Action Space
GPAgentBench-2K is introduced, the first Constrained MDP (CMDP) LLM-agent benchmark for primary-care clinical decision-making, constructed from expert-validated records of real-world GP encounters, and uncovers a clinical quality-safety gap.
Bo-Qi Chen, Xudong Liu, Y. Ao et al.
· 0 citations