Skip to content
Open access

HyBioSum: A Hybrid Framework for Biomedical Long-Document Summarization With Supervised Extractive and LLM-Based Generation

2026 · IEEE Access · Vol 14, pp. 136245-136266 · 0 citations · 89 references

TL;DR

A hybrid extract-then-summarize framework that first identifies salient sentences using a supervised extractive model and then generates an abstractive summary through LLM prompting is proposed, which improves efficiency by focusing the generative process on informative content while reducing the processing cost typically associated with long biomedical documents.

Abstract

Biomedical text summarization (BTS) aims to automatically condense single or multiple biomedical documents into concise summaries while preserving essential information. However, biomedical texts are often lengthy, structurally complex, and rich in domain-specific terminology, making automated summarization particularly challenging. Although recent Large Language Models (LLMs) have demonstrated strong performance across various NLP tasks, their direct application to biomedical documents remains difficult due to computational constraints and the risk of losing clinically important information in long inputs. To address these limitations, we propose a hybrid extract-then-summarize framework that first identifies salient sentences using a supervised extractive model and then generates an abstractive summary through LLM prompting. This strategy improves efficiency by focusing the generative process on informative content while reducing the processing cost typically associated with long biomedical documents. We evaluate the proposed approach on benchmark datasets, including PubMed and CORD-19, using ROUGE, METEOR, and BERTScore metrics. Additionally, we conduct a human evaluation assessing relevance, conciseness, informativeness, and readability. Experimental results show that our method achieves competitive performance compared to recent state-of-the-art models. Ablation studies further confirm the benefits of integrating supervised extraction with LLM-based generation within a unified hybrid framework.

Read PDF

Similar papers

Open access Sep 2026

Evaluating Large Language Models for Biomedical Text Summarization: A Study of Cardiovascular Research

In this study, presented a comprehensive evaluation of abstractive and extractive summarization performance across three prominent large language models (LLMs): ChatGPT, DeepSeek, and Gemini. A total of 8,000 cardiovascular-related research abstracts were collected from PubMed and summarized using two distinct promptin...

Burcu Baştürk, Aytuğ Onan · 0 citations
#natural language process... Preprint Sep 2026

A Hybrid Hierarchical 1D-CNN-BiLSTM Framework for Extractive Summarization of Biomedical and Clinical Text

Large language models have made abstractive summarization remarkably fluent, but generated summaries can hallucinate facts, posing serious risks in biomedical and clinical domains. We address this by removing generation from the pipeline and framing summarization as extractive sentence selection. Our Hybrid Hierarchica...

Saad Bin Ather, Muhammad Saif, A. H. Khan et al. · 0 citations
Open access 2026

Knowledge Distillation for Biomedical Text Classification: A Systematic Comparative Analysis of Multiple Teacher–Student Architectures

Findings demonstrate that compact models can achieve strong biomedical classification performance through KD under compatible teacher–student pairings, while also highlighting that KD effectiveness varies substantially depending on the specific model combination.

Amine Gonca Toprak, Aytuğ Onan · 0 citations
Open access Sep 2026

Systematic assessment of text summarization methods for biomedical literature from frequency methods to language models

Summary The rapid expansion of biomedical literature demands automated summarization tools that reliably condense research articles into concise, accurate summaries. We benchmarked 62 summarization methods, ranging from frequency-based and TextRank extractors to encoder-decoder models (EDMs) and large language models (...

Fabio Baumgärtel, Enrico Bono, Lucas Fillinger et al. · 0 citations
2026

Assessing Transformer Models for Abstractive Summarization of Scientific Articles

The results show that BART achieves the best performance with an ROUGE-2 F1-score of 0.40664, while T5 demonstrates superior grammatical acceptability, achieving 93.36%, but BART achieves a very near performance to T5.

Emad Nabil · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.