Skip to content
Open access

Variantscape: Large Language Model-Driven Mining of Biomedical Literature for Clinical Interpretation of Cancer Variants

Aug 2026 · medRxiv · 0 citations
Medicine

TL;DR

Variantscape has the potential to support MTB workflows and translational research by rapidly revealing signals from underlying abstracts, and offers a practical resource for accelerating discovery and supporting precision oncology research and translation.

Abstract

Background: Precision oncology relies on accurate interpretation of tumour-detected gene variants, to guide personalized treatment decisions. However, accurate interpretation of variants in context requires extensive information that is often buried within unstructured biomedical literature and obscured by inconsistent nomenclature, making manual retrieval labour-intensive and prone to omissions. Methods: To address this challenge, we developed Variantscape, a large-scale, automated pipeline and open-access web tool. It integrates traditional natural language processing methods with state-of-the-art large language models to extract, standardize, and analyze co-associations between genetic variants, cancer types, and therapeutic interventions from published biomedical abstracts. Findings: From over 3 million abstracts screened, 335,817 gene name-containing articles were eligible for downstream extraction. Among these, 7,423 (2.2%) simultaneously mentioned a variant, cancer type, and therapeutic agent, encompassing 3,902 unique variants across 98 cancer types and 388 therapeutic agents. This highlights the inefficiency of manual literature retrieval in molecular tumour board (MTB) workflows. Network analysis revealed 14,831 statistically significant co-associations, represented in a literature-derived graph with 4,388 nodes and 46,943 edges. Canonical alterations in well-studied cancers (e.g., BRAF V600E in melanoma) were strongly linked to established treatments, while several rare variants also emerged with high-confidence literature support. Interpretation: By applying large language models to biomedical literature, Variantscape enables scalable, context-aware extraction of trilateral variant-treatment-cancer relationships. This approach supports early evidence synthesis/hypothesis generation, highlights underrecognized or rare associations, and offers a practical resource for accelerating discovery and supporting precision oncology research and translation. Unlike static databases, Variantscape is continuously updatable and leverages large language model-based inference to uncover putative associations without manual curation. Variantscape has the potential to support MTB workflows and translational research by rapidly revealing signals from underlying abstracts.

Read PDF

Similar papers

Open access Aug 2026

megaMine: a scalable, rule-based framework for mining gene-cancer-drug evidence from biomedical literature

The rapid expansion of the oncology literature has outpaced manual curation of clinically relevant gene-cancer-drug associations and oncogenic driver evidence. Existing automated approaches often lack transparency or are difficult to scale across heterogeneous data sources. To address this gap, we developed megaMine, a transparent, rule-based, and context-aware literature-mining framework that integrates therapeutic and driver evidence from PubMed, PubTator, and Europe PMC by combining entity recognition, hierarchical heuristics, and contextual labeling. In therapy mode, megaMine was applied to approximately 100,000 oncology articles published between 2015 and 2025, yielding more than 23,000 structured sentence-level evidence records, with standardized annotations for drug response, resistance, and study context. Internal evaluation of context labels showed the strong separability between efficacy and non-efficacy evidence using ridge logistic regression (AUROC = 0.915; AUPRC = 0.941). Benchmarking against NCI/OncoKB-supported drug-cancer associations showed that curated clinical associations had higher megaMine composite evidence scores than unlabeled comparison pairs [median (IQR): 25.6 (9.07-72.5) vs. 3.61 (1.69-8.69); Wilcoxon rank-sum test, P < 2.2 × 10−16]. In driver mode, megaMine retrieved mutation- and biomarker-related evidence from an ERBB-focused gastric cancer query, generating 750 evidence rows from 200 PMIDs. These results demonstrate that deterministic and interpretable approaches can support scalable evidence extraction for downstream applications such as knowledge graph construction and literature-based evidence synthesis.

Muhammad Junaid, K. Prazanowska, Ha-Eun Jeong et al. · 0 citations
Review Open access Aug 2026

OncoGenRAG: Evidence-Grounded Retrieval and BioBERT Classification for Precision Oncology Variant Interpretation

The increasing use of tumor sequencing has intensified the need for fast, traceable interpretation of genomic variants. General-purpose large language models can produce fluent answers, but unsupported statements, weak provenance, and stale knowledge limit their suitability for clinical genomics. We developed OncoGenRAG, a research framework that combines a parameter-efficiently fine-tuned BioBERT classifier with an entity-aware retrieval system over a curated, multi-source oncology knowledge base. The reported knowledge base contains 933 harmonized records derived from CIViC, ClinVar/dbSNP, Open Targets, UniProtKB/Swiss-Prot, Ensembl Variation, and linked PubMed literature. The classifier assigns one of five labels: Pathogenic, Likely Pathogenic, Variant of Uncertain Significance, Benign, or Oncogenic; the retrieval component ranks evidence records using subword TF-IDF similarity and explicit gene, variant, and cancer-type matches. A rejection rule suppresses answers when retrieval support is below a prespecified threshold. In the authors’ held-out evaluation, the classifier achieved 92.40% accuracy, 93.15% weighted precision, 92.40% weighted recall, and 92.65% weighted F1 score. In a separate benchmark of 100 clinical-style queries, OncoGenRAG achieved reported Precision@1 of 94.5%, Precision@3 of 96.8%, and 100% database grounding. No hallucinated answer was observed under the study’s operational definition, compared with a 41.0% no-hallucination rate for the ungrounded baseline. These results should be interpreted as internal validation rather than proof of universal safety because query construction, annotator agreement, class-specific performance, calibration, and external validation data were not available for independent analysis. OncoGenRAG provides a transparent design for evidence retrieval and abstention, but it is a research prototype and must not be used to select treatment without expert review.

Amaan Arif, José Valentim dos Santos · 0 citations
Review Open access Aug 2026

A machine learning framework for predictive interpretation of variants of uncertain significance in hereditary cancer

Introduction Variant interpretation remains a major bottleneck in clinical genomics, with variants of uncertain significance (VUS) representing a critical unresolved challenge due to insufficient evidence for definitive classification. Existing in silico tools exhibit variable and often inconsistent performance complicating clinical decision-making, particularly in the context of hereditary cancer genomics. Methods In this study, we developed a machine learning framework trained on 1,04,646 high-confidence ClinVar germline variants (3-star+ review status) annotated with Ensembl VEP (v114, GRCh38) and CADD v1.6 pathogenicity scores to classify variants as Pathogenic or Benign, subsequently applying the trained model to reclassify 40894 ClinVar VUS. Train/test partitioning was performed at the variant level (80/20 split) to prevent data leakage, with hyperparameter optimization via GridSearchCV and performance assessed by 10-fold cross-validation. Four classifiers were evaluated viz. Logistic Regression, Support Vector Machine, Random Forest and XGBoost, with Random Forest achieving the highest performance (AUC-ROC = 0.9995, 95% CI: 0.9993–0.9997; 10-fold CV AUC = 0.9992 ± 0.0004). Probability thresholds of P ≥ 0.80 (Pathogenic) and P <= 0.20 (Benign) were derived from Precision-Recall curve analysis, achieving empirically validated precision of 99.63% and 99.77% respectively on held-out test variants. Results and Discussion Applied to 40,894 ClinVar VUS, the model reclassified 19393 (47.4%) as Likely Pathogenic and 8,957 (21.9%) as Likely Benign, while 12,544 (30.7%) were conservatively retained as uncertain. External validation on 7,462 ENIGMA-classified BRCA1/BRCA2 variants from the BRCA Exchange database, completely independent of the ClinVar training data demonstrated an overall concordance of 98.83% (AUC = 1.0000). Further validation of VUS reclassification against 671 variants classified as VUS in ClinVar but definitively classified by ENIGMA yielded an overall concordance of 89.57% (Pathogenic: 96.4%, Benign: 87.4%). SHAP-based explainability analysis confirmed that predictions were predominantly driven by biologically interpretable features, including CADD Phred score, VEP functional impact tier, variant consequence class and population allele frequency, consistent with ACMG/AMP evidence criteria. This reproducible pipeline provides a clinically grounded computational approach to VUS triaging in precision oncology, with external validation supporting its generalizability to independent hereditary cancer gene datasets.

Nayeema Nizamuddin, Soham Biswas, Akshaykumar Zawar et al. · 0 citations
Review Open access Aug 2026

The Role of Advanced Language Models in MedicalDiagnostics: A Case Study on Breast Cancer Prediction

Breast cancer remains one of the most common and life threatening cancers worldwide, and early detection is strongly associated with improved survival and reduced treatment burden. This study  investigates the ability of Large Language Models to perform diagnostic prediction from structured breast cancer related data. We systematically evaluated 12 LLMs across three public datasets with different clinical characteristics: the Wisconsin Breast Cancer Dataset (WBCD) based on cytological features, the Breast Cancer Coimbra Dataset (BCCD) based on metabolic biomarkers, and the Mammographic Mass Dataset (MMD) based on mammographic attributes. The evaluation covered 13 prompting strategies, including three zero-shot variants, few-shot prompting, and multiple Chain-of-Thought (CoT) and knowledge-enhanced reasoning settings. Performance was assessed using confusion-matrix-based metrics, including accuracy, precision, recall, F1-score, specificity, and Matthews Correlation Coefficient. The results showed that performance was strongly dependent on both dataset type and prompting design, and no single model dominated all tasks. The best model–strategy pairvaried by dataset: Cogito-v1-preview-qwen-32B achieved the highest F1-score on WBCD with an F1-score of 92.00%, GPT 4.1 and GPT 4o on BCCD with an F1-score of 85.39%, and Gemini 2.5 Flash Lite on MMD with an F1-score of 82.91%. Prompt engineering had a substantial effect on outcomes, but its benefit varied across models, with some systems improving under knowledge-enhanced few shot prompting and others performing best under simpler strategies. Comparison with state-of-the-art traditional ML baselines showed that, while LLMs do not yet surpass supervised methods, the performance gap has narrowed substantially, particularly on MMD, where the best single-run gap was 2.04 percentage points in F1 and the mean gap under robustness analysis was approximately 5.3 points. Robustness analysis across multiple few-shot example sets confirmed stable performance on WBCD (F1 = 91.37 ± 0.71%) and BCCD (F1 = 87.46 ± 1.91%), while revealing moderate sensitivity on MMD (F1 = 79.62 ± 2.87%). Although the evaluated LLMs did not outperform traditional supervised models, the study provides a clear performance baseline for future research on structured clinical prediction with language models. The results show that LLMs may offer value as complementary exploratory tools, but their outputs should be interpreted only with expert oversight because clinically significant errors remain.

Habibe Karayiğit, F. Kalelioğlu · 0 citations
Open access Aug 2026

OncoRAG: graph-based retrieval enabling clinical phenotyping from oncology notes using local mid-size language models

Oncology notes contain the richest clinical detail, yet they remain largely inaccessible at scale because extracting structured phenotypes requires either substantial language model infrastructure, curated training data, or cloud computing under regulatory constraints. We developed OncoRAG, combining ontology enrichment, knowledge graph construction, graph-diffusion reranking, and structured prompting with a locally deployed 14B-parameter language model without model weight fine-tuning. Applied to three cohorts—triple-negative breast cancer (TNBC; 104 patients, 42 features; primary development), recurrent high-grade glioma (RiCi; 191 patients, 19 features; cross-lingual and cross-disease evaluation with cohort-specific configuration), and MIMIC-IV (100 patients, 10 features; limited external evaluation on overlapping features)—OncoRAG achieved F1 scores of 0.80, 0.79, and 0.84, improving over direct large language model (LLM) prompting and naive retrieval-augmented generation (RAG) baselines by 0.19–0.22 and 0.17–0.19 F1, and outperforming direct prompting with a 5× larger 70B model by 0.09–0.10 F1. In an exploratory survival analysis (12 events), both feature sets showed close point estimates of the C-index (0.77 vs 0.76), but equivalence cannot be statistically confirmed given the limited event count. OncoRAG enables accurate clinical phenotyping from multilingual oncology notes using a locally deployable mid-size model, without model weight fine-tuning or external data sharing.

P. Salome, Maximilian Knoll, David Walz et al. · 0 citations
Open access Aug 2026

A Deterministic Framework for Integrated Genome Variant Interpretation - The ‘GenomeVAP’

High-throughput genomic sequencing generates vast amounts of data, yet the interpretation of individual-genetic variants remains hindered by the dispersion of relevant evidence across various databases. We present a modular, web-based framework (GenomeVAP) designed for deterministic evidence integration in genomic research. Unlike machine learning models that often introduce noise into genomic annotations or rely on opaque predictive thresholds, GenomeVAP utilizes a weighted, rule-based scoring methodology to synthesize evidence from primary repositories, including ClinVar,[1] dbSNP,[2] Ensembl,[3] and the GWAS Catalog.[4] We evaluate the framework's efficacy through representative batch entries of clinically significant variants, demonstrating that centralized, automated retrieval reduces manual querying time while maintaining high transparency and reproducibility. GenomeVAP is intended exclusively as a bioinformatics software framework to aid interpretation in healthcare and academic research fields

Unknown authors · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.