Skip to content
Open access

Artificial Intelligence in Plastic Surgery Education: Insights from Parallel Turkish and English Versions of a Board Examination

Jul 2026 · Applied Sciences · 0 citations · 30 references

Abstract

Background: Large language models (LLMs) increasingly pass medical board examinations and aid clinical knowledge retrieval and decision support; validating their specialist knowledge is a prerequisite for safe use. Two limitations weaken existing evidence: many studies reuse public questions that may be in training data, and most evaluations use only English-language examinations. Plastic surgery serves here as a representative, highly specialized subfield of medicine. Methods: We used the non-public board examination of the Turkish Society of Plastic, Reconstructive, and Aesthetic Surgeons (TSPRAS): 100 single-best-answer (SBA) items in official parallel Turkish and English versions. Six contemporary LLMs from three developers, in matched free and paid tiers, completed six runs per language (7200 responses) under consistent English instruction. Accuracy, inter-run reliability, and cross-language and tier differences were analyzed. Results: All models exceeded the 60% passing threshold (71.5–86.4% aggregate), with almost perfect inter-run agreement (Cohen’s κ ≥ 0.843). Question language produced no significant difference after correction, and item difficulty correlated strongly across languages (Pearson r = 0.928); divergent items reflected content, not language. Paid variants offered only incremental advantages. Conclusions: Current LLMs showed no significant accuracy difference between the Turkish and English examination versions under a constant English prompt, supporting supervised study-aid use on this MCQ format.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.