Skip to content
Open access

Topic Modeling and Semantic Similarity-Based Evaluation of Biomedical Text Abstracts Generated by LLM

Jul 2026 · Bilişim Teknolojileri Dergisi · Vol 19, pp. 287-298 · 0 citations · 16 references

TL;DR

The results indicate that abstractive summaries, particularly those generated by Gemini and ChatGPT, achieve coherence levels comparable to or higher than original abstracts under LDA and LSA, while extractive summaries exhibit greater variability.

Abstract

This study investigates how large language models (LLMs)—ChatGPT, Gemini, and DeepSeek—preserve thematic consistency in biomedical text summarization. A dataset of 56,000 cardiovascular texts was constructed, including 8,000 original PubMed abstracts and 48,000 LLM-generated summaries. After domain-specific preprocessing, topic modeling was performed using Latent Dirichlet Allocation (LDA), Latent Semantic Analysis (LSA), and Non-negative Matrix Factorization (NMF) across topic sizes (k = 10–50). Thematic consistency was evaluated using C_v coherence scores. In addition, semantic similarity between original and generated texts was assessed using SBERT-based cosine similarity to capture meaning preservation beyond lexical overlap. The results indicate that abstractive summaries, particularly those generated by Gemini and ChatGPT, achieve coherence levels comparable to or higher than original abstracts under LDA and LSA, while extractive summaries exhibit greater variability. For NMF, original abstracts show more stable trends, whereas ChatGPT’s extractive summaries remain competitive at higher topic sizes. Overall, coherence tends to decrease with increasing topic numbers in LDA, whereas LSA and NMF benefit from larger topic sizes. These findings highlight the importance of combining coherence and semantic similarity for a more comprehensive evaluation of LLM-based biomedical summarization.

Read PDF

Similar papers

Open access Sep 2026

Evaluating Large Language Models for Biomedical Text Summarization: A Study of Cardiovascular Research

In this study, presented a comprehensive evaluation of abstractive and extractive summarization performance across three prominent large language models (LLMs): ChatGPT, DeepSeek, and Gemini. A total of 8,000 cardiovascular-related research abstracts were collected from PubMed and summarized using two distinct promptin...

Burcu Baştürk, Aytuğ Onan · 0 citations
Open access Sep 2026

Systematic assessment of text summarization methods for biomedical literature from frequency methods to language models

Summary The rapid expansion of biomedical literature demands automated summarization tools that reliably condense research articles into concise, accurate summaries. We benchmarked 62 summarization methods, ranging from frequency-based and TextRank extractors to encoder-decoder models (EDMs) and large language models (...

Fabio Baumgärtel, Enrico Bono, Lucas Fillinger et al. · 0 citations
Open access 2026

Enhancing Biomedical Multi-Label Text Classification via Topic-Based Text Representation

: Biomedical texts naturally contain multiple biological and medical concepts within a document, resulting in a semantically rich and complex structure. Consequently, multi-label text classification (MLTC) has become a suitable framework for comprehensively modeling biomedical texts, including clinical reports, laborat...

Oyku Berfin Mercan, Nezihe Turhan Turan, Aytuğ Onan · 0 citations
Open access 2026

HyBioSum: A Hybrid Framework for Biomedical Long-Document Summarization With Supervised Extractive and LLM-Based Generation

A hybrid extract-then-summarize framework that first identifies salient sentences using a supervised extractive model and then generates an abstractive summary through LLM prompting is proposed, which improves efficiency by focusing the generative process on informative content while reducing the processing cost typica...

Azzedine Aftiss, Salima Lamsiyah, Christoph Schommer et al. · 0 citations
Open access Aug 2026

Benchmarking MeSH-augmented embeddings for biomedical document similarity

The extensive volume of biomedical scientific literature requires efficient methods for retrieving relevant documents based on semantic technologies and biomedical concepts. While embedding-based methods have shown improvements over traditional keyword-based methods, the integration of domain-specific terminologies lik...

Rohitha Ravinder, Lukas Geist, Nelson Quiñones et al. · 0 citations
Open access 2026

Knowledge Distillation for Biomedical Text Classification: A Systematic Comparative Analysis of Multiple Teacher–Student Architectures

Findings demonstrate that compact models can achieve strong biomedical classification performance through KD under compatible teacher–student pairings, while also highlighting that KD effectiveness varies substantially depending on the specific model combination.

Amine Gonca Toprak, Aytuğ Onan · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.