Skip to content
Open access

Optimizing Multilingual Embedding Models for Retrieval and Reranking in RAG Pipelines: Enhancing Semantic Search in Turkish Medical Datasets

Aug 2026 · SN Computer Science · Vol 7 · 0 citations · 33 references

Abstract

This study examines the effectiveness of enhanced multilingual embedding models in improving retrieval performance for Turkish medical text data. We consider two specific medical applications in Turkish language: the TUS examination, a standardized medical assessment featuring exam questions, and Clinical QA, which involves authentic patient-physician interactions. By implementing multi-stage fine-tuning protocols on domain-specialized models, we provide detailed performance assessment and explore cross-domain transfer capabilities of the trained models. Our findings indicate that domain-specialized models improve in-domain retrieval relative to generic models, and that systematic optimization through our multi-stage pipeline yields measurable gains in retrieval precision. For instance, domain-specific fine-tuning improves TUS retrieval performance from 0.69 to 0.79 in P@1 and from 0.77 to 0.85 in MRR, while Clinical QA fine-tuning with hard-negative sampling improves P@1 from 0.33 to 0.39 and MRR from 0.41 to 0.48 relative to the vanilla encoder. In addition, reranking improves P@1 from 0.788 to 0.823 in our evaluated setting, corresponding to a 4.5% relative improvement. Furthermore, we find that, in multilingual model training, domain-specific knowledge acquired in the healthcare context of one language effectively transfers and enhances performance across other languages.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.