The Alignment Paradox of Medical Large Language Models in Infertility Care: Decoupling Algorithmic Improvement From Clinical Decision-Making Quality
Abstract Background Large language models (LLMs) have been proposed as decision-support tools in assisted reproductive technology (ART), but it remains unclear whether posttraining alignment strategies translate into clinically acceptable decision support. Outcome-based benchmarks may reward token-level correctness whi...