Skip to content
Review Open access

Context-Grounded PET-CT Report Generation Using LoRA-Fine-Tuned BioMedLM with Physician Validation and Safety Evaluation

Sep 2026 · Indian Journal of Nuclear Medicine · Vol 41, pp. 461-480 · 0 citations · 50 references

TL;DR

Parameter-efficient fine-tuning of biomedical language models can generate coherent physician-style PET-CT report sections from structured clinical findings, with high physician-assessed factual correctness and no unsafe hallucinations in this limited reviewed subset, supporting the feasibility of structured-finding-to-narrative report generation for clinical AI pipelines.

Abstract

Positron emission tomography-computed tomography (PET-CT) reporting is cognitively demanding, requiring integration of quantitative metabolic data and anatomical findings across multiple body regions. Existing segmentation and analysis models provide anatomical and physiological parameters, but not cohesive physician-style reports; however, they can supply a structured clinical context for a dedicated PET-CT language model. Existing large language model (LLM) approaches for radiology report generation predominantly target chest radiography and lack domain-specific grounding for nuclear medicine, and unconstrained free-text generation from generic prompts introduces clinically unacceptable hallucination risk. We assembled 836 section-level PET-CT reports from 170 patients (85 breast cancer, 85 controls) across four primary anatomical regions (Brain excluded due to stereotyped language) plus a whole-body report for each patient. The biomedical language model BioMedLM (2.7B) was fine-tuned using Low-Rank Adaptation (LoRA) with a faithfulness-augmented objective penalising hallucination and rewarding negation preservation. Structured clinical context, including maximum standardised uptake value (SUVmax), laterality, and lesion localisation, served as grounded input. Evaluation included five-fold patient-level cross-validation, held-out test performance, temporal and cross-group validation, clinical safety metrics, and blinded independent review by two nuclear medicine physicians across 28 test cases. Physician review showed factual correctness 4.59/5, completeness 4.77/5, and clinical usefulness 4.12/5, with no unsafe hallucinations identified (one-sided 95% CI: 0-10.7%). Clinical acceptability (Grade A) was 53.6% (Physician 1) and 57.1% (Physician 2; combined 55.4%), with moderate inter-rater agreement (κ = 0.61). BioMedLM-LoRA achieved Recall-Oriented Understudy for Gisting Evaluation (ROUGE-L) 0.545 and Bidirectional Encoder Representations from Transformers Score (BERTScore-F1) 0.911 on the held-out test set, outperforming template and retrieval-augmented baselines (all p < 10 -10 ); SUVmax faithfulness was 1.000 and negative-finding preservation 0.790. Term frequency-inverse document frequency (TFIDF) retrieval achieved higher ROUGE-L (0.814), reflecting lexical-overlap advantage in this closed, single-site corpus. Automated laterality assessment remained proxy-level (accuracy 0.271), underscoring rule-based evaluation limitations. Not all flagged laterality (33/120) and hallucination-proxy (14/120) cases underwent physician adjudication; these findings represent supportive evidence from a limited subset, not a generalisable safety claim. Parameter-efficient fine-tuning of biomedical language models can generate coherent physician-style PET-CT report sections from structured clinical findings, with high physician-assessed factual correctness and no unsafe hallucinations in this limited reviewed subset, supporting the feasibility of structured-finding-to-narrative report generation for clinical AI pipelines. However, the structured findings used as model inputs were derived from reference reports rather than independently generated by an upstream image-analysis system, representing a limitation. Larger external, multi-institutional studies using independently extracted image-derived context are required to establish generalisability and assess suitability for clinical deployment under physician supervision.

Read PDF

Similar papers

Preprint Sep 2026

MetaStructAtlas: A Grounded 3D Vision-Language Dataset and Benchmark for Functional and Structural Reasoning in Whole-Body PET/CT

The joint interpretation of metabolic function and anatomical structure is essential for clinical diagnosis in whole-body PET/CT. Although recent advances in 3D medical vision-language models have demonstrated remarkable progress, current efforts are limited to regional CT imaging, leaving a critical void in comprehens...

Chen-Guang Zheng, Le Xue, Yi-Chi Zhang et al. · 0 citations
Open access Oct 2026

UniPET: a unified approach for whole-body CT to PET translation

Positron emission tomography (PET) provides critical metabolic information for oncological imaging, yet its use is constrained by radiation exposure, cost, and limited availability. Synthesizing PET-like images from computed tomography (CT) has been proposed as a way to approximate metabolic information; however, exi...

Francesco Di Feola, V. Guarrasi, Mikael Johansson et al. · 0 citations
Open access Aug 2026

Transformative Medical Report Generation for Brain MRI using BERT-GPT4 Hybrid Models and Real-Time Classification

The need for automated medical report generation is critical in healthcare, especially in the radiology domain where generating structured, coherent, and clinically accurate brain MRI reports remains labor-intensive in process. Most of the available methods cannot contextually interpret sparse keyword input or accurate...

Wani H. Bisen, Avinash J. Agrawal · 0 citations
Sep 2026

A Resource-Efficient End-to-End Multimodal AI Framework for Automated PET-CT Report Generation and TNM Staging in Breast Cancer

Positron Emission Tomography-Computed Tomography (PET-CT) is integral to breast cancer staging. However, manual interpretation is time-consuming and subject to inter-observer variability. While artificial intelligence has shown promise in lesion detection and standardised uptake value (SUVmax) quantification, exist...

D. Prakash, Venkatesh Rangarajan, Manoj Kumar · 0 citations
#computer vision Review Sep 2026

NV-Reason-CT: 3D Visual Language Model for CT Analysis

The NV-Reason-CT model, a generative vision--language model for chest and abdominal CT combining native 3D visual encoding with radiologist-guided reasoning, and the model and training code are released to support reproducible research on explainable AI for volumetric medical imaging.

Andriy Myronenko, Dong Yang, Yu-Cheng Tang et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Error Detection for PET/CT Radiology Reports: Domain-Specific vs Large Language Models

It is suggested that domain-specific training matters more than model scale for PET/CT report error detection, supporting compact models as an accurate and computationally efficient approach to automated radiology report quality assurance.

Hermione Warr, Harry Anthony, Lilli J. Freischem et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.