Skip to content
Conference

Deep Learning for Cancer Type Identification from RNA-Seq Gene Expression: A Comparison with ML

Aug 2026 · 2026 International Conference on Modern Sustainable Systems (CMSS) · pp. 1404-1407 · 0 citations · 21 references

Abstract

Deep learning has changed how images and text are analyzed, and it is now widely used on gene expression data as well. This paper asks a direct question. On a standard cancer classification task from RNA-Seq expression, does deep learning actually beat simple machine learning, and when does the difference show up. We use the public TCGA pan-cancer RNA-Seq dataset, with 801 samples across five tumor types and more than 20000 genes. We build three neural models, a multilayer perceptron, a one-dimensional convolutional network, and an autoencoder paired with a softmax classifier, and we compare them with a random forest and a linear support vector machine. Every model is tested with stratified 5-fold cross-validation repeated over three seeds. On the full data the models are close. The multilayer perceptron reaches a macro F1 of 0.9996, matching the linear support vector machine and slightly ahead of the random forest, while the convolutional network trails at 0.920 because the gene axis has no spatial order for a convolution to use. The clearer story appears when labels are scarce. With only about 28 labeled samples the multilayer perceptron leads the random forest by roughly two points, and the gap closes as data grows. The lesson is that a plain neural network is a strong and honest baseline here, that convolution needs a matching structure to help, and that the value of deep learning on this task is largest when labels are few.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.