Skip to content
Review Open access

Text2FHIRwallet: Automated Generation of FHIR Patient Summaries from Unstructured Cardiology Reports Using Fine-Tuned Portuguese Language Models—Development and Evaluation of a Health Professional Wallet

Aug 2026 · Applied Sciences · Vol 16, pp. 7795 · 0 citations · 24 references

TL;DR

Text2FHIRwallet demonstrates that domain-specific fine-tuning of a Portuguese-language pretrained language model achieves near-ceiling clinical NER accuracy, enabling scalable, interoperable, and privacy-compliant patient summary generation from unstructured cardiology text, offering an end-to-end pathway for integrating AI-driven NLP into clinical workflows and FHIR-based health information ecosystems.

Abstract

Background and Objectives: Cardiology departments generate large volumes of unstructured free-text reports that impose substantial manual review burdens on clinicians; at Hospital de Santa Maria—Portugal’s largest public hospital—manual review of 12,651 reports took approximately seven minutes per report, representing over 1475 h of avoidable administrative work. This study presents Text2FHIRwallet, a health professional digital wallet that automates extraction and structuring of clinical entities from unstructured Portuguese cardiology reports using fine-tuned Named Entity Recognition (NER) models and maps the results to Fast Healthcare Interoperability Resources (FHIR) R4 patient summaries. Materials and Methods: Following the Design Science Research Methodology (DSRM) and CRISP-DM, we fine-tuned four transformer-based models—BERTimbau Base, BERTimbau Large, Albertina PT-PT, and MediAlbertina—on 305 manually annotated cardiology reports (77,309 tokens; κ = 0.85 inter-annotator agreement) covering eight clinical entity types, drawn from a corpus of 12,651 anonymised documents. Entities were mapped to FHIR R4 resources and delivered through a secure, role-based mobile wallet (React Native). Evaluation comprised token-level NER benchmarking with bootstrapped confidence intervals and McNemar’s testing, FHIR mapping accuracy assessment on 100 manually reviewed reports, processing-efficiency measurement, and a usability pilot with 10 cardiologists (SUS, NPS). Results: MediAlbertina achieved the highest NER performance (macro F1 = 0.985, 95% CI: 0.979–0.990), significantly outperforming all baseline models (p < 0.01, McNemar’s test) and comparing favourably with—though not directly comparable to, given differing languages and datasets—published benchmarks such as GPT-4 (F1 = 0.962 in ophthalmology NER) and fine-tuned BERT models for lung cancer NER (F1 ≈ 0.85–0.90). FHIR mapping accuracy was 98% on 100 independently reviewed reports. Report processing time was reduced from approximately seven minutes to 15–30 s (93–96% reduction), with peak batch-inference throughput of up to 1000 reports/h under parallelised GPU load (observed end-to-end throughput in pilot deployment was approximately 250 reports/h). The pilot usability evaluation yielded a SUS score of 87 (excellent) and an NPS of 80. Conclusions: Text2FHIRwallet demonstrates that domain-specific fine-tuning of a Portuguese-language pretrained language model achieves near-ceiling clinical NER accuracy, enabling scalable, interoperable, and privacy-compliant patient summary generation from unstructured cardiology text, offering an end-to-end pathway for integrating AI-driven NLP into clinical workflows and FHIR-based health information ecosystems, with implications for administrative efficiency, care coordination, and clinical research in non-English-language settings.

Read PDF

Similar papers

Review Aug 2026

From text to insight: a systematic literature review of keyword and keyphrase extraction techniques for healthcare text mining

This review provides researchers and practitioners with a structured framework for method selection based on their specific constraints and identifies six prioritized research directions for future investigation, identifying critical research gaps including the preservation of multi-word clinical concepts, scarce evaluation in real-world clinical workflows, and persistent hallucination risks in model-based extraction.

Mouhamed Gaith Ayadi · 0 citations
Conference Jul 2026

Generating Reliable Synthetic Clinical Discharge Summaries for Medical Text Analysis

Clinical text is an important part of healthcare systems because it is used to store and manage patient information in documents such as discharge summaries, doctor notes, and diagnostic reports. Among these documents, discharge summaries are especially important because they provide a brief overview of a patient’s diagnosis, treatment procedures, medications, and follow-up instructions after hospitalization. These summaries are also useful for healthcare research and medical data analysis. However, strict privacy regulations and hospital policies restrict access to real clinical records, making it difficult for researchers to collect large datasets for developing and testing machine learning models in healthcare. To address this issue, this study proposes a framework for generating and validating synthetic clinical discharge summaries using transformer-based biomedical language models. Initially, the clinical text is preprocessed using cleaning, formatting, and tokenization techniques to improve consistency and readability. Biomedical language models such as BioBERT, RoBERTa, and DistilBERT are then used to generate contextual embeddings and capture semantic relationships within medical text. In addition, semantic similarity analysis, entailment-based validation, and faithfulness evaluation are applied to verify the consistency and reliability of the generated summaries while preserving patient privacy and maintaining clinical relevance.

Mohammad Imran, M. Irfan, Sajida Sultana.Sk et al. · 0 citations
Open access Aug 2026

Developing an open-source framework for LLM evaluation of patients using EHR clinical documentation; performance of LLMs relative to medical professionals

Current LLMs do not achieve inter-rater reliability levels comparable to medical professionals in clinical information extraction from ENT documentation, suggesting they are best suited for initial extraction with human verification rather than autonomous operation.

L. Barrett, N. Joshi, A. S. North et al. · 0 citations
Review Open access Aug 2026

NERFlow: A Workflow-Based Subsystem of FIT4NER for LLM-Assisted Medical Named Entity Recognition

Preparing training data for domain-specific medical Named Entity Recognition (NER) involves a trade-off between annotation quality, expert effort, and data privacy: manual annotation is costly, whereas cloud-based Large Language Models (LLMs) raise concerns about the control of sensitive clinical text. This article introduces NERFlow, a workflow-driven subsystem of the FIT4NER project whose contribution is an abstraction layer that makes rule-based, model-based, and LLM-based annotation interchangeable and comparable within one configurable workflow environment. Open-source LLMs are integrated as exchangeable annotation services, deployable locally or in cloud-agnostic infrastructures via Kubernetes, and embedded into the KM-EP knowledge management system. NERFlow was evaluated qualitatively, through a cognitive walkthrough, an IEEE 1028 technical review, and a user-centered survey with 18 participants, and quantitatively on the CRAFT corpus with seven open-source and hosted LLMs run through an identical pipeline. The results support its use as LLM-assisted pre-annotation with expert correction, with locally deployable open-source models as the more reliable basis for reproducible operation.

Florian Freund, Philippe Tamla, Bao Tran et al. · 0 citations
Open access Feb 2025

Enhancing Large Language Models for Identifying and Prioritizing Important Medical Jargons From Electronic Health Record Notes Using Data Augmentation: Comparative Study

This study evaluated both closed-source and open-source large language models for extracting and prioritizing medical jargon from EHR notes relevant to individual patients, leveraging prompting techniques, fine-tuning, and data augmentation and found that model performance could deviate largely based on prompting styles.

W. Jang, Sharmin Sultana, Zonghai Yao et al. · 1 citation
Open access Jul 2026

Benchmarking large language models for clinical data extraction from Portuguese medical notes in a university hospital

The findings support the feasibility of applying LLM-based natural language processing tools in resource-limited, non-English healthcare settings and should assess emerging high-parameter models and explore additional clinical domains.

Breno Gabriel Araújo Sampaio de Jesus, Tomaz Castrillon Figueiredo, Clariele de Almeida Pereira et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.