Skip to content

Author

Kang-Su Ha

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Review Open access Jul 2026

A Structural Labeling-File Audit and Secondary Model-Output Evaluation of Korean Specialized and Essential Medical Knowledge Datasets for Medical AI Research

Background: Korean-language medical question-answering datasets are increasingly used for large language model (LLM) development, but structural completeness alone does not establish model performance or clinical validity. We examined two national Korean medical knowledge datasets by combining a public labeling-file audit with a secondary evaluation of item-level outputs distributed with the corresponding official LLM packages. Methods: We reviewed the official documentation and complete public download inventories, parsed 34 training/validation labeling archives containing 31,036 records, and audited required fields, duplication, text length, question type, and specialty distribution. We also audited the official Qwen2.5-14B LoRA model packages and independently recalculated performance from their distributed item-level KorMedMCQA output files (2494 identical items per model). Accuracy was reported with Wilson 95% confidence intervals; model outputs were compared using an exact McNemar test and a paired bootstrap confidence interval. Because no compatible local GPU was available, model inference was not independently rerun. Results: The documentation described 34,487 labeled QA pairs, of which 31,036 (90.0%) were present in the publicly accessible training/validation labeling files; no public test-labeling archive was listed. Required fields were complete, no duplicated qa_id values were found, and one duplicated question-answer pair occurred in the Essential dataset. Multiple-choice items comprised 78.7% of all documented QA pairs, and the top three domains comprised 56.8%. In the distributed KorMedMCQA outputs, the Essential-care model answered 1603/2494 items correctly (64.27%; Wilson 95% CI 62.37–66.13%), while the Specialized-medicine model answered 1596/2494 correctly (63.99%; 95% CI 62.09–65.85%). The paired difference was 0.28 percentage points (bootstrap 95% CI −0.44 to 1.00), with no significant difference by exact McNemar test (46 vs. 39 discordant correct items; p = 0.515). Conclusions: The public training/validation labeling files showed favorable basic structural completeness, but the unavailable test labels, multiple-choice predominance, domain imbalance, limited record-level provenance, and incomplete model-package traceability constrain claims of clinical readiness. The distributed model outputs demonstrated moderate examination-style benchmark performance without a significant difference between the two models. These findings support use as research infrastructure, not evidence of clinical validity, and indicate the need for independent inference reproduction, clinician-led open-ended and safety evaluation, temporal updating, specialty-stratified reporting, and human oversight before clinical use.

Mi-Ae Yang, Kang-Su Ha · 0 citations
Review Open access Aug 2026

A Human-Governed Clinical Informatics Framework for Safe AI-Assisted Mental Health Counseling: Secondary Framework Development and Requirement Mapping Study.

BACKGROUND Natural language processing and large language model systems are increasingly used to support mental health documentation, screening, and follow-up planning. In counseling contexts, model outputs may influence diagnostic framing, risk recognition, and clinical record content. Static performance metrics and fluent generated summaries are not sufficient to support safe implementation without governance, safety gating, human review, and monitoring. OBJECTIVE This study aimed to develop a human-governed clinical informatics framework for safe AI-assisted mental health counseling and make the formative evidence base and requirement-mapping process traceable. METHODS We conducted a secondary framework development and requirement mapping study using the Korean AI Hub psychological counseling dataset, official data description and use documents, released KLUE-BERT risk prediction model materials, released KoAlpaca summary generation resources, and a deidentified 139-case rule-based summary safety screening audit table derived from the original summary comparison file. Raw counseling transcript text, reference summary full text, and generated summary full text are not included in the manuscript or supplementary materials. We extracted failure modes from documented data and model characteristics, released code and configuration files, documentation-reported model metrics, and rule-based proxy flags. Each failure mode was mapped to safety controls, operational criteria, and deployment-level requirements. RESULTS The official documents described 1661 counseling sessions and 465,474 paragraph-level tokens across depression, anxiety disorder, addiction, and normal control groups. Of the 1661 sessions, the documented split included 1339 (80.6%) training, 173 (10.4%) validation, and 149 (9%) test sessions. The summary generation materials documented 1278 training summaries and 139 test summaries. Documentation-reported model metrics included KLUE-BERT accuracies of 71.43% for depression, 73.53% for anxiety, and 66.67% for addiction and KoAlpaca BERTScore precision, recall, and F1-score values of 62.13%, 59.56%, and 60.80%, respectively. The 139-case screening table contained 77 (55.4%) depression, 31 (22.3%) anxiety, and 31 (22.3%) addiction cases. Rule trigger rates included unsupported content proxy flags in 41% (57/139) of cases, overdiagnostic expression proxy flags in 31.7% (44/139) of cases, medicalized expression proxy flags in 54.7% (76/139) of cases, and any rule-based proxy flag in 91.4% (127/139) of cases. These values are conservative rule trigger rates rather than confirmed clinical error rates. The findings informed a 7-stage workflow, 6 safety control layers, an operational safety gate, a workflow-to-control crosswalk, deployment-level transition criteria, and a constructed high-risk example. CONCLUSIONS AI-assisted mental health counseling should be implemented as a governed clinical information workflow rather than as an autonomous diagnostic or documentation pathway. The proposed framework specifies safeguards and validation requirements for future supervised evaluations, but it does not itself establish clinical safety or clinical effectiveness. Prospective simulation, clinician usability testing, patient or client feedback, and independent expert validation remain necessary before routine deployment.

Mi-Ae Yang, Kang-Su Ha · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.