Recall Is Not Execution: a Large Language Model and the Revised Geneva Score [Dataset].
Study summary A frontier large language model (Claude Opus 4.8, model identifier claude-opus-4-8, Anthropic) was tested against the revised Geneva score for pulmonary embolism using simulated patient profiles. Two tasks were run. Task 1, weight elicitation. The model was asked to assign numeric weights to the component...