LifeSciBench: Evaluating Language Models on Realistic, Expert-Level Tasks in the Life Sciences
LifeSciBench is introduced, a benchmark of 750 expert-authored tasks designed to evaluate whether language models can handle realistic life science research work, with each constituent task paired with a human expert-written rubric.
Amelia Liu, Andrew Ho, Anne Marie Droste et al.
· bioRxiv · 2 citations