Skip to content

A medical retrieval augmented generation prototype for somali symptom guidance with cross lingual evidence retrieval and technical verification

Sep 2026 · Discover Artificial Intelligence · Vol 6 · 0 citations · 15 references
Topic Modeling

TL;DR

This work asks whether an implemented cross-lingual retrieval-augmented generation prototype can accept Somali symptom narratives, retrieve English medical-textbook evidence, and return a structured Somali response with explicit safety boundaries and page-level citations.

Abstract

Somali remains underrepresented in evaluated health-oriented natural-language systems. Direct use of a general-purpose large language model for symptom questions is difficult to audit because answers may not be traceable to a defined medical source and multilingual performance may differ substantially from English. This work asks whether an implemented cross-lingual retrieval-augmented generation prototype can accept Somali symptom narratives, retrieve English medical-textbook evidence, and return a structured Somali response with explicit safety boundaries and page-level citations. Five English medical textbooks were extracted and cleaned, including OCR correction for CURRENT Medical Diagnosis & Treatment. Text was divided into overlapping 275-word chunks with 40-word overlap. A total of 64,177 chunks were embedded with OpenAI text-embedding-3-small (1536 dimensions) and indexed using FAISS IndexFlatIP over L2-normalized vectors. At runtime, a Somali query is translated into clinical English, checked by deterministic bilingual emergency rules, embedded, and matched using cosine similarity (top-k = 5; minimum score = 0.25). Retrieved passages are supplied to OpenAI GPT-4o-mini (API identifier: gpt-4o-mini) under an evidence-only prompt. A citation guard rejects unavailable or missing chunk identifiers, and the structured answer is translated back to Somali without changing urgency. The prototype uses Flutter and FastAPI. Deterministic software verification covered data preprocessing, document cleaning, chunk generation, embedding consistency, vector-index construction, retrieval execution, metadata alignment, citation validation, deterministic safety routing, API validation, and structured response generation. The defined verification scenarios completed successfully, and the final index contains 64,177 vectors with book, section, and page metadata. An illustrative end-to-end transaction confirmed that the implemented components could execute as an integrated pipeline. These findings establish functional software implementation only; no clinical accuracy, linguistic quality, user acceptance, model superiority, or deployment effectiveness is claimed. The contribution is a reproducible prototype architecture that separates cross-lingual translation, deterministic emergency escalation, semantic retrieval, evidence-constrained generation, citation validation, and Somali presentation. The system has not undergone clinician-labelled triage evaluation or prospective clinical testing and must not be used for diagnosis, prescribing, or autonomous emergency decisions.

Read PDF

Similar papers

Conference Aug 2026

RAG-Powered Clinical Conversational Agent for Guideline Retrieval: A Prototype System Comparing Standard and HyDE-Enhanced Retrieval

Healthcare professionals spend significant time manually retrieving information from clinical guidelines during time-sensitive scenarios. This study presents a prototype conversational AI agent that leverages Retrieval-Augmented Generation (RAG) to enable rapid, natural-language access to obstetric clinical guidelines....

Isabel Sofía Tovar Sánchez, Rubén Manrique, Nathalia Ortega et al. · 0 citations
#large language models Review Open access Sep 2026

Large Language Models for Clinical Note Simplification: A Systematic Review and Experimental Evaluation of Medical Text Readability.

The findings suggest that conventional readability metrics should be extended with domain-specific measures to more accurately assess comprehensibility in medical texts and that large Language Models show strong potential to enhance the accessibility of clinical documentation for patients.

M. Teichmann, Pelin Özkara Menekseoglu, Julian Schwarz et al. · 0 citations
#large language models Open access Sep 2026

Medical Concept Normalization of German Clinical Expressions to SNOMED CT.

INTRODUCTION Clinical narratives in electronic health records frequently contain clinical expressions describing medical conditions. Their free-text format limits interoperability and automated processing. Medical concept normalization (MCN) addresses this challenge by mapping textual expressions to standardized termin...

Helena Adam, Akhila Abdulnazar, Roland Roller et al. · 0 citations
Open access Sep 2026

Hybrid lexical-semantic retrieval over SNOMED CT: combining two retrieval paradigms to facilitate clinical data entry

Objectives: Searching a reference terminology such as SNOMED CT confronts two retrieval technologies of fundamentally opposite nature. Deterministic lexical matching is precise, order-independent and character-level, and lets a clinician refine a query incrementally, but it is confined to the wording and language of th...

A. L. Osornio, A. Hoejen, K. Kewley · 0 citations
Review Open access Aug 2026

Evaluating Clinical Concept Extraction and Evidence-Bounded Terminology Linking: Multisite Model Comparison and Pilot Ablation Study

Measured extraction performance varied substantially by matching definition, whereas exact-link decisions varied with the availability of matched terminology evidence, support separate evaluation of extraction, retrieval, evidence-grounded linking, and extension candidacy.

Y. Chen, M. Popescu · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.