Skip to content
Conference

Hybrid AI Framework for Multi-Omics-based Kidney Tumor Subtype Classification and Precision Oncology

Jul 2026 · 2026 4th International Conference on Sustainable Computing and Smart Systems (ICSCSS) · pp. 1637-1643 · 0 citations · 15 references

Abstract

This study presents an AI-driven multi-omics framework designed to improve the classification of kidney tumor subtypes and support personalized treatment strategies in precision oncology. By integrating genomic, transcriptomic, and proteomic data from publicly available TCGA repositories, the proposed model builds comprehensive molecular profiles of kidney tumors. The system achieves a classification accuracy exceeding 96% by combining machine learning and deep learning techniques — specifically Random Forest (RF) for feature selection, Support Vector Machines (SVM) for handling high-dimensional data, Convolutional Neural Networks (CNNs) for spatial pattern extraction, and Transformer models to capture contextual relationships across biologically ordered gene sequences. Unlike conventional ensemble approaches, this hybrid framework is optimized for both predictive accuracy and computational efficiency, making it suitable for real-time clinical use. Additionally, the model identifies key tumor-specific biomarkers that can guide individualized therapy. Experimental results confirm that the proposed system significantly outperforms traditional diagnostic methods and single-omics models. Designed for integration with cloud-based Clinical Decision Support Systems (CDSS), the framework has strong potential to enhance real-time oncology decision-making.

View source

Similar papers

Open access Aug 2026

A heterogeneous multimodal ensemble framework for multi-omics breast cancer prognosis

Accurate breast cancer prognosis remains a major challenge in precision oncology due to tumor heterogeneity and the complexity of integrating high-dimensional multi-omics data. Although multimodal learning approaches have improved predictive performance by combining clinical and molecular information, many existing methods rely on a single ensemble strategy that remains susceptible to prediction variance and limited robustness in high-dimensional, low-sample-size biomedical datasets. This study investigated whether integrating complementary ensemble strategies within a unified multimodal framework could improve the robustness and predictive performance of breast cancer prognosis. A heterogeneous multimodal ensemble framework was developed in which stacking was used to integrate complementary information from clinical, gene expression, and copy number variation (CNV) data through meta-learning, while bagging was incorporated to stabilize the meta-learning process via bootstrap aggregation. The outputs of the stacking and bagging branches were combined using weighted probability fusion. The framework was evaluated on the METABRIC breast cancer cohort and compared with unimodal models and a conventional stacking ensemble using an independent test set and stratified tenfold cross-validation. The proposed hybrid framework achieved a ROC-AUC of 0.936, outperforming unimodal clinical and molecular models (ROC-AUC = 0.8140.885) and the conventional stacking ensemble (ROC-AUC = 0.898). Stratified tenfold cross-validation further demonstrated consistent improvements in mean ROC-AUC, recall, F1-score, balanced accuracy, and Matthews correlation coefficient, indicating improved robustness and stable performance across the internal validation folds. On the independent test set, the hybrid framework reduced false-negative predictions and increased sensitivity relative to the stacking ensemble, demonstrating a more favorable balance between identifying high-risk patients and maintaining overall predictive performance. Rather than introducing a new ensemble algorithm, this study demonstrates that assigning complementary roles to stacking multimodal information integration and bagging for prediction stabilization provides an effective and robust framework for multi-omics breast cancer prognosis. The proposed hybrid strategy consistently improved predictive performance and robustness compared with conventional stacking while demonstrating stable performance across internal validation, supporting the use of complementary ensemble paradigms for multimodal prediction in precision oncology.

Reza Bozorgpour, Mohammadreza Soltany Sadrabadi · 0 citations
Review Jul 2026

AI-Driven Multi-Omics Integrated Applications Using Diverse Neural Networks for Breast Cancer Diagnostic Screening and Biomarker Discovery.

Breast cancer (BC) is a highly complex and heterogeneous malignancy and the most prevalent cancer among women worldwide. The diagnosis, prognosis, and the treatment of BC pose significant challenges that are responsible for their limited therapeutic efficacy. Omics-based technologies have gained substantial attention in BC diagnosis through molecular profiling and diverse clinical analytics. The integration of metabolomics, proteomics, transcriptomics, and genomics provides a multidimensional approach to personalized BC diagnosis and treatment through high-throughput molecular profiling. Moreover, the emergence of artificial intelligence (AI) has also supported more accurate and early diagnosis of BC through multimodal integration of diverse datasets. The integration of advanced deep learning (DL) and machine learning (ML) has been extensively exploited for tumor grading, histopathological classification, molecular profiling, diagnostic imaging, and prognostic prediction. This review aims to summarize recent developments in AI-driven multi-omics approaches for the discovery of BC biomarkers. We have also highlighted the integration of omics-based data like metabolomics, proteomics, transcriptomics, and genomics with key AI techniques, including ML and DL, that play a crucial role in the inclusion of multi-omics in cancer and biomarker discovery. We have further discussed AI-based BC screening and diagnostic approaches, as well as the contribution of AI models for patient stratification, biomarker discovery, and prediction of therapeutic response. Additionally, key limitations and challenges, including data heterogeneity, high computational complexity, and model interpretability, have also been highlighted in the present review. Conclusively, we have also outlined future perspectives on the integration of AI and multi-omics to revolutionize precision clinical medicine and improve clinical outcomes in BC theranostics.

V. Kumari, Harshita Tiwari, Swati Singh et al. · 1 citation
Open access Sep 2026

A Systematic Machine Learning Framework for Evaluating and Ranking Omics Layers in Cancer Drug Response Prediction

This study introduces a top-down framework for evaluating the utility of multi-omics features to predict the response of 309 drugs in cancer cell lines. This was done by taking a multi-omics approach where data from proteomic, transcriptomic, genomic, metabolomic, and miRNA were integrated with drug sensitivity (area under the curve, AUC) data. We performed modular dimensionality reduction using t-SNE (t-distributed Stochastic Neighbor Embedding), followed by K-Means clustering to stratify cell lines into data-driven molecular subgroups, and applied a Random Forest model to refine the drug list, selecting only those with a prediction accuracy exceeding 75%. Our findings show that among the evaluated single-omics features, transcriptomics is the most informative; however, multi-omics integration significantly enhances predictive capability compared to single-omics analysis, with a combination of transcriptomic, proteomic, and miRNA data achieving the best predictive performance across both primary and validation datasets. Cluster analysis showed the importance of well-defined clusters, indicating that while silhouette scores were linked to prediction success, biological variability also played a critical role. This study advances personalized oncology treatment strategies and provides a foundation for future studies focused on ranking omics features based on their predictive capabilities, eventually contributing to better therapeutic outcomes. Predictive performance is used here to evaluate omics feature strength, rather than as an objective to optimize predictive models.

Unknown authors · 0 citations
Jul 2026

Attention-enhanced Multi-omics Model for Pan-cancer Drug Response Prediction and Biomarker Discovery

Cancer drug discovery remains challenged by tumour heterogeneity and limited experimental scalability. Recently, virtual drug screening integrating machine learning algorithms has yielded numerous research results, some of which have been successfully patented and are expected to be further translated and deployed in drug discovery pipelines. Most current models for virtual drug screening fail to effectively integrate multi-omics data or capture nonlinear cross-omics interactions, restricting predictive accuracy and biomarker discovery across diverse cancers with high heterogeneity. We developed a multi-omics fusion deep learning model integrating mutation, methylation, transcriptomic, and metabolomic profiles from over 900 pan-cancer cell lines derived from the DepMap database. Our framework synergizes random forest-based feature selection to prioritize biologically relevant omics features and multi-head attention mechanisms to model nonlinear interactions between cellular multi-omics landscapes. Further biological analysis of the selected features enabled the deciphering of potential biomarkers related to drug effects. Our multi-omics fusion model attained high-performance drug response prediction across nearly 900 cancer cell lines (median Pearson r = 0.50 vs Pearson r = 0.22 for the former model for all included drugs). Validation demonstrated robust accuracy for the MEK inhibitor Trametinib (r = 0.78, MAE = 0.59) and the non-oncology agent BMOV (r = 0.73). The model identified BRAF-mutant melanoma sensitivity and PI3K/AKT bypass resistance, consistent with existing findings. Feature mining revealed TERT modulation and oxidative stress induction as BMOV's probable anticancer mechanisms, while Benzamide targeted metabolic vulnerabilities. This study introduced a deep-learning-based multi-omics fusion model to predict pan-cancer drug response. The framework achieved the expansion of the anticancer spectrum of existing anticancer drugs and explored the potential anticancer effects of non-anticancer drugs. Moreover, the integration of SHAP/MDI feature interpretation algorithms enabled mechanistic biomarker discovery. However, the prediction results that were not reported in previous research are yet to require further experimental verification. In conclusion, this work established a robust DL-driven platform for virtual drug screening and biomarker discovery, providing a computational platform that could aid virtual drug screening and biomarker discovery and facilitate the development of precision oncology.

Yi-Bo Wang, Xin-Wen Zhang, Yi-Dan Ye et al. · 0 citations
Open access Aug 2026

A hybrid ensemble-based parallel learning framework for multi-omics data integration and cancer subtype classification

Integrating multi-omics data to understand biological processes in human diseases is a complex bioinformatic task. Machine learning (ML), particularly deep learning (DL) models, offers a promising approach to multi-omics data integration and analysis. However, existing DL models generally integrate multi-omics data by concatenating the input data space or learned feature space, which is a sub-optimal approach. In addition, single classifiers are commonly used in DL-based methods, which can compromise the performance. Furthermore, the gradient descent optimization technique in DL suffers from a high computational cost and local sub-optimal solutions. To address these challenges, this article presents a novel cancer subtype classification framework using multi-omics integration and an ensemble-based parallel DL/ML architecture. Specifically, a multimodal autoencoder is used for effective feature learning across omics types, overcoming the limitations of naïve concatenation. A hybrid ensemble model comprising DL and ML learners with a meta-learner enhances classification robustness beyond single models. To improve optimization and computation, we incorporate a hybrid Back-Propagation and Particle Swarm Optimization (PSO) strategy and execute the entire framework on a parallel processing platform, reducing computation time while enhancing global search capability. The proposed framework is evaluated empirically with two benchmark data sets from The Cancer Genome Atlas (TCGA), namely the TCGA Pan-cancer and TCGA Breast Invasive Carcinoma (BRCA) data sets. The results indicate a high performance with accuracy rates of 89.51% and 90.9% for TCGA Pan-cancer and TCGA BRCA, respectively. The parallel implementation of the proposed framework reduces the computation time, resulting in a speed-up of 3 times and 2.5 times for TCGA Pan-cancer and TCGA BRCA, respectively. The findings ascertain the efficacy of the proposed framework for the classification of cancer subtypes, offering a promising solution for implementation in real-world environments.

Mohammed Nasser Al-Andoli, Shing Chiang Tan, Kok Swee Sim et al. · 0 citations
Open access Jul 2026

EMMA-STRAT: a multi-omics based machine learning framework for stratification of endometrial carcinoma molecular subtypes and MSI status

Uterine Corpus Endometrial Carcinoma (UCEC) is the most common gynecologic malignancy, with molecular heterogeneity influencing prognosis and treatment response. Although TCGA-defined molecular subtypes and multi-omics datasets have improved biological understanding of UCEC, externally evaluated computational frameworks for molecular stratification remain limited. To address this, we developed EMMA-STRAT, a supervised multi-omics machine learning framework integrating mRNA expression, miRNA expression, and DNA methylation data to classify UCEC genomic subtypes and microsatellite instability (MSI) status. Using the TCGA cohort (N = 433) for model development and internal validation, we benchmarked six classifiers and evaluated final model performance on two independent Clinical Proteomic Tumor Analysis Consortium (CPTAC) cohorts (N = 95 and N = 108). Multi-omics integration consistently outperformed single-omics models, with RNA expression as the strongest standalone modality. For MSI-H versus MSS classification, a LightGBM model trained on 20 SVM-selected features per omics layer achieved an internal balanced accuracy of 98.1% and external balanced accuracies of 93.1–94.9%. For four-class genomic subtyping, a Multi-Layer Perceptron trained on 50 LASSO-selected features per omics layer achieved an internal balanced accuracy of 89.1% and external balanced accuracies of 84.7–86.2%. Both models showed favorable discrimination and probability calibration relative to reference baselines, although calibration estimates for low-prevalence classes including POLE should be interpreted cautiously. SHapley Additive exPlanations (SHAP)-based interpretability analysis identified model-selected features including MLH1, CDKN2A, PPP4R4, and hsa-miR-378a, with downstream analyses supporting their biological plausibility. All results are openly accessible via an interactive browser at https://naisarg14.github.io/EMMA-STRAT-web-viewer/index.html. EMMA-STRAT provides an externally evaluated, research-grade computational framework for multi-omics molecular stratification of endometrial carcinoma. Integration of mRNA, miRNA, and DNA methylation data supported prediction of MSI-H versus MSS status and TCGA-defined genomic subtypes across independent cohorts. However, since EMMA-STRAT requires multi-omics data and was not directly compared with established clinical classifiers, it should currently be interpreted as a research-oriented molecular stratification framework rather than a clinically deployable decision-making model. The developed framework provides a basis for future prospective validation, incorporation of clinicopathological variables, and direct comparison with ProMisE-based or integrated clinical risk models.

Naisarg Patel, A. Salumets, V. Modhukur · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.