Jul 2026· Modern Pathology· pp.
101043
· 0 citations· 50 references
Medicine
TL;DR
Overall, the evidence suggests that success in lower capability levels does not consistently translate to higher-level reasoning, especially when reports are inconsistent, required staging inputs are missing, or clinical assumptions must be inferred, which helps explain gaps between benchmark results and practical adoption.
Abstract
Pathology reports anchor cancer diagnosis and staging, yet their narrative structure limits reliable translation into structured, machine-actionable knowledge, creating a bottleneck between expert interpretation and scalable clinical intelligence. Despite decades of clinical natural language processing (NLP) research, pathology text remains among the most complex and consequential sources of medical data to operationalize at scale. Large language models (LLMs) offer new approaches for reading, extracting, and interpreting these reports. We synthesize current LLM work in cancer pathology using a four-level capability framework across the pathology report data lifecycle: (level 1) text preparation and quality checks, (level 2) information extraction, (level 3) guideline-based clinical reasoning, such as TNM staging and registry coding, and (level 4) interpretive synthesis, such as explanations, summarization, or decision support. Rather than grouping studies by NLP task labels, this framework tracks how LLM applications progress from preprocessing and extraction toward higher-level interpretation and synthesis. We followed PRISMA-ScR guidelines and searched four databases through September 2, 2025, identifying 41 eligible studies. Most studies focus on level 2 tasks, with fewer addressing level 3 and level 4 tasks. Encoder-based models, including domain-specific variants such as BioBERT, were commonly used for structured extraction tasks, whereas generative models, including GPT, LLaMA, and Mistral-family models, were increasingly evaluated for prompting-based extraction, staging, and summarization. Reported performance was often high for well-defined extraction tasks, but external validation was uncommon, and metrics varied across studies, limiting direct comparison. Overall, the evidence suggests that success in lower capability levels does not consistently translate to higher-level reasoning, especially when reports are inconsistent, required staging inputs are missing, or clinical assumptions must be inferred, which helps explain gaps between benchmark results and practical adoption. Future work should prioritize robust multi-site validation, clinically meaningful error analysis, transparent evaluation, and privacy-preserving implementation strategies to support safe integration in oncology.
This is the first study to benchmark SLMs on Italian EHRs and investigate the role of clinical expertise in prompt engineering, offering valuable insights for the future integration of SLMs into real-world clinical workflows.
Federica Corso, V. Peppoloni, L. Mazzeo et al.· Communications Medicine· 0 citations
PURPOSE
To develop a large language model (LLM) (Truveta Language Model Oncology [TLM-Oncology]) to extract real-world oncology staging data across multiple cancer types from clinical documentation with high precision.
METHODS
We selected patients from a large integrated health system with a bladder, cervical, colorectal, breast, or prostate cancer diagnosis in their structured data. We identified relevant notes using note metadata and keywords and annotated overall stage; T, N, and M; associated timeframe; and cancer diagnosis on a sample of 700 notes as ground truth. Of the 700 notes, 450 were divided equally between training, validation, and test sets for bladder, cervical, and colorectal cancers; 150 were used for targeted error-pattern training on these cancers; and the remaining 100 were split equally between breast and prostate cancer test sets. We started with a pretrained LLM and applied supervised fine-tuning to adapt the model to structured clinical information extraction. Model performance was measured using precision, recall, and F1 scores at the relation level and individual attribute level.
RESULTS
We extracted over 2.5 million staging records for 217,768 patients from over two million notes. Relation-level precision across the six attributes ranged from 0.77 to 1.0 for the first three cancers and, without further training, 0.83 to 1.0 for two additional cancers.
CONCLUSION
TLM-Oncology extracted detailed cancer staging information for five cancers from a variety of clinical documentation within a single integrated health system with high precision and turned data that were previously inaccessible into a valuable resource for downstream use. We are currently evaluating TLM-Oncology on other solid tumors within three additional health systems to assess its generalizability.
S. Abhyankar, Rajesh Rao, Mehraveh Salehi et al.· JCO Clinical Cancer Informat...· 0 citations
Structured reporting converts free-text radiology narratives into queryable data keys, facilitating cohort assembly, longitudinal tracking, and training label generation for medical AI. The prevailing paradigm follows a two-stage pipeline: (1) constructing a reporting template, (2) extracting information to populate it. While the extraction stage has benefited from advances in large language models (LLMs), template construction remains a manual bottleneck relying on labor-intensive expert consensus that is static, difficult to scale, and may fail to capture real-world reporting diversity. We address this limitation with ASTAR, an LLM-based framework for Automated induction of STAndardized radiology Reporting templates from large-scale clinical free-text corpora. Extensive experiments on 4,215 fetal brain MRI reports from multiple centers demonstrate that, in this reporting scenario, the ASTAR-induced template surpasses two expert-curated templates across template coverage, information fidelity, diagnostic fidelity, and expert-rated usability, reducing template development from weeks of committee deliberation to hours of automated processing.
This paper summarizes the main challenges currently facing, including medical data privacy and labeling problems, interpretability and clinical credibility barriers of the model, and systemic barriers to multimodal fusion.
This study investigates whether large language models (LLMs) can assist in identifying extraction errors and generating candidate rules to improve symbolic clinical NLP systems and suggests that LLMs can support scalable rule refinement for symbolic clinical NLP systems.
N. Wang, A. Kakadiaris, C. Li et al.· medRxiv· 0 citations
Large language models (LLMs) show promise for text-based pathology tasks, yet most reported applications remain experimental, lack formal clinical validation, or operate outside secure, health system-approved environments. We developed and clinically validated a rule-guided, agent-based LLM that assists gastrointestinal (GI) biopsy reporting by automating report structuring while preserving full diagnostic authority with the pathologist. The AI agent (Microsoft 365 Copilot) ran within an enterprise-approved, HIPAA-compliant Microsoft 365 environment, configured with a fixed rule-based system configuration prompt and a quick-text knowledge base. In a prospective validation, 94 GI biopsy cases were evaluated by subspecialty GI pathologists using specimen container labels extracted from the laboratory information system and pathologist-entered shorthand diagnoses. Agent outputs were reviewed for formatting accuracy, organ and procedure identification, shorthand expansion fidelity, blank diagnosis enforcement, and diagnostic safety. The agent preserved specimen part structure and correctly identified organ, sub-organ, and procedure context in 100% of cases; shorthand expansion was accurate in all applicable cases. Minor formatting deviations occurred in 8 cases (8.5%) without affecting diagnostic meaning. Two cases (2%) showed minor diagnostic misinterpretation, in which descriptive container-label terms (e.g., "ulcer," "erosion") were incorporated into diagnostic text; no hallucinated diagnoses were identified. Repeatability testing on cases enriched for descriptive labels showed 81% identical outputs across 105 runs (19% variability), with non-reproducible semantic leakage in 3% of runs. A comparative time study showed faster AI-assisted reporting (mean 39 vs 72 seconds for speech-to-text and 76 seconds for manual typing; ∼33-37 second reductions, p < 0.05), measured across the full workflow through sign-out, with lower variability. By restricting this end-to-end, production-embedded agent to rule-guided structuring, formatting, and controlled shorthand expansion while prohibiting diagnostic inference, the system achieved high efficiency, consistency, and seamless workflow integration on real GI biopsy cases. Low-frequency, stochastic errors and minor variability remain inherent to LLMs despite strict constraints; although infrequent, they indicate such systems are best suited for non-diagnostic, clerical augmentation rather than autonomous use. All output therefore requires pathologist careful review before sign-out. These findings support constrained, agent-based LLMs to safely enhance reporting efficiency while preserving diagnostic responsibility and human oversight.
Ibrahim Abukhiran, Akila Mansour, M. L. Caicedo et al.· Modern Pathology· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.