Author

Aashamsu Nepal

1 paper indexed here

Fetches their full publication history.

Not the right person? Other researchers publish under this name.

Book Open access Jun 2026

When Medical LLMs Eat Their Own Output: The Effects of Recursive Self-Training on German Medical Text

Since the widespread adoption of large language models (LLMs), AI-generated text has become increasingly prevalent in scientific communication and online content. As the proportion of LLM generated text grows, concerns have emerged regarding recursive self-training, where models are trained on data generated by earlier model versions. Prior work suggests that such training regimes can lead to model collapse, characterized by the loss of semantic diversity and degradation of learned data distributions. Although these effects have been studied primarily in general-domain English settings, their implications for domain-specific and high-stakes biomedical applications remain insufficiently understood. We investigate the impact of recursive self-training on German-language medical text by recursively fine-tuning a medically specialized language model over multiple iterations. For comparison purposes we present two alternative training pipelines. In the first pipeline, the model is recursively trained exclusively on synthetic German medical text generated by earlier model versions. In the second pipeline, synthetic German medical text is combined with human German medical text at each iteration to test whether data blending mitigates recursive degradation. The evaluation is based on answer accuracy achieved on a set of multiple-choice questions (MCQs) drawn from the German medical licensing examination, the Staatsexamen, as well as lexical diversity metrics such as distinct-n, and next-token probability concentration measures that capture distributional collapse. Across multiple training iterations, the purely synthetic pipeline exhibits progressive degradation in both question–answering performance and linguistic diversity, consistent with the effects of model collapse. In contrast, although minor degradation is still observable, the hybrid pipeline maintains substantially more stable performance over time. These results suggest that human-anchored recursive training (blending synthetic data with real, human-generated data) constitutes a promising mitigation strategy against recursive degradation.

Aashamsu Nepal, Ruben Nuredini, Gerrit Meixner · 0 citations