A hybrid extract-then-summarize framework that first identifies salient sentences using a supervised extractive model and then generates an abstractive summary through LLM prompting is proposed, which improves efficiency by focusing the generative process on informative content while reducing the processing cost typically associated with long biomedical documents.
Abstract
Biomedical text summarization (BTS) aims to automatically condense single or multiple biomedical documents into concise summaries while preserving essential information. However, biomedical texts are often lengthy, structurally complex, and rich in domain-specific terminology, making automated summarization particularly challenging. Although recent Large Language Models (LLMs) have demonstrated strong performance across various NLP tasks, their direct application to biomedical documents remains difficult due to computational constraints and the risk of losing clinically important information in long inputs. To address these limitations, we propose a hybrid extract-then-summarize framework that first identifies salient sentences using a supervised extractive model and then generates an abstractive summary through LLM prompting. This strategy improves efficiency by focusing the generative process on informative content while reducing the processing cost typically associated with long biomedical documents. We evaluate the proposed approach on benchmark datasets, including PubMed and CORD-19, using ROUGE, METEOR, and BERTScore metrics. Additionally, we conduct a human evaluation assessing relevance, conciseness, informativeness, and readability. Experimental results show that our method achieves competitive performance compared to recent state-of-the-art models. Ablation studies further confirm the benefits of integrating supervised extraction with LLM-based generation within a unified hybrid framework.
In this study, presented a comprehensive evaluation of abstractive and extractive summarization performance across three prominent large language models (LLMs): ChatGPT, DeepSeek, and Gemini. A total of 8,000 cardiovascular-related research abstracts were collected from PubMed and summarized using two distinct promptin...
Burcu Baştürk, Aytuğ Onan· Sakarya University Journal o...· 0 citations
Large language models have made abstractive summarization remarkably fluent, but generated summaries can hallucinate facts, posing serious risks in biomedical and clinical domains. We address this by removing generation from the pipeline and framing summarization as extractive sentence selection. Our Hybrid Hierarchica...
Saad Bin Ather, Muhammad Saif, A. H. Khan et al.· 0 citations
Findings demonstrate that compact models can achieve strong biomedical classification performance through KD under compatible teacher–student pairings, while also highlighting that KD effectiveness varies substantially depending on the specific model combination.
Summary The rapid expansion of biomedical literature demands automated summarization tools that reliably condense research articles into concise, accurate summaries. We benchmarked 62 summarization methods, ranging from frequency-based and TextRank extractors to encoder-decoder models (EDMs) and large language models (...
Fabio Baumgärtel, Enrico Bono, Lucas Fillinger et al.· iScience· 0 citations
The results show that BART achieves the best performance with an ROUGE-2 F1-score of 0.40664, while T5 demonstrates superior grammatical acceptability, achieving 93.36%, but BART achieves a very near performance to T5.
Emad Nabil· Islamic University Journal o...· 0 citations