Making medical AI benchmarks clinically interpretable: the case of mental health
Abstract Medical artificial intelligence (AI) benchmarks are increasingly used to assess the readiness of large language models for health-related tasks, but aggregate performance scores can obscure clinically meaningful variation across domains. HealthBench, an open benchmark of 5000 multi-turn health conversations, r...