Skip to content
#protein folding Open access

CRISPR dependency profiles distinguish pediatric from adult cancer cell lines via explainable machine learning: Identification of GPX4 and proliferative genes as therapeutic targets

Oct 2026 · Intelligent Data Analysis · 0 citations · 39 references
Cancer Genomics and Diagnostics CRISPR and Genetic Engineering

Abstract

Pediatric and adult cancers diverge profoundly in their molecular architecture, yet whether these differences extend to genome-scale genetic dependencies remains poorly characterized, with prior studies examining single subtypes or individual algorithms rather than a systematic, interpretable comparison. To address this gap, we tested whether Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR)-Cas9 loss-of-function dependency profiles can discriminate pediatric from adult cancer cell lines, and applied explainable artificial intelligence (XAI) to identify the genes driving this distinction. Chronos-normalized CRISPR GeneEffect scores for 18,531 genes across 1208 cell lines (242 pediatric, 966 adult) were analyzed; class imbalance was corrected with the Synthetic Minority Over-sampling Technique (SMOTE) within each training fold, differentially dependent genes were identified by Mann-Whitney U tests with Benjamini-Hochberg False Discovery Rate (FDR) correction, and 200 features were retained by SelectKBest. Eleven classifiers were benchmarked under 5-fold cross-validation across 11 metrics, and predictions were interpreted with SHapley Additive exPlanations (SHAP) and Local Interpretable Model-agnostic Explanations (LIME) In this first (to our knowledge) genome-scale, multi-classifier comparison Light Gradient Boosting Machine (LightGBM) achieved the best discrimination (Area Under the Receiver Operating Characteristic Curve [ROC-AUC] = 0.909, accuracy = 88.7%, Matthews Correlation Coefficient [MCC] = 0.632, Brier score = 0.089), with the leading ensemble and margin-based models reaching ROC-AUC values near 0.90. Among 3602 differentially dependent genes (FDR < 0.05), Integrin-Linked Kinase (ILK), Talin-1 (TLN1), and Aurora Kinase A (AURKA) showed greater dependency in pediatric lines, whereas Glutathione Peroxidase 4 (GPX4) and Insulin-like Growth Factor 2 messenger RNA-Binding Protein 1 (IGF2BP1) were comparatively more essential in adult lines. SHAP highlighted CHMP4B, IGF2BP1, GPX4, and LDB1, and the representative true-positive LIME case emphasized IGF2BP1, PAQR6, and CDS2. By exposing dependencies that adult-centric precision-oncology programs overlook, this work provides a framework for age-specific target discovery and prioritizes the pediatric-enriched ILK–TLN1 mechanotransduction axis for experimental validation.

Read PDF

Similar papers

#computer vision Review Open access May 2015

A survey study on major technical barriers affecting the decision to adopt cloud services

The comparison of adopter and non-adopter sample reveals three potential adoption inhibitor, security, data privacy, and portability, which underlines the importance of the technical and security perspectives for research investigating the adoption of technology.

Nattakarn Phaphoom, Xiaofeng Wang, S. Samuel et al. · 111 citations · ⚡8
#computer vision Open access Feb 2018

Lean Internal Startups for Software Product Innovation in Large Companies: Enablers and Inhibitors

This study investigates how Lean internal startup facilitates software product innovation in large companies and identifies its enablers and inhibitors, and shows the potential of the method-in-action framework to investigate the Lean startup approach in non-startup context.

Henry Edison, Nina M. Smørsgård, Xiaofeng Wang et al. · 78 citations · ⚡6
#computer vision Book Open access Jul 2015

Understanding the affect of developers: theoretical background and guidelines for psychoempirical software engineering

This paper highlights the challenges to conduct proper affect-related studies with psychology, provides a comprehensive literature review in affect theory, and proposes guidelines for conducting psychoempirical software engineering.

D. Graziotin, Xiaofeng Wang, P. Abrahamsson · 56 citations · ⚡4
#machine learning Open access May 2017

What Influences the Speed of Prototyping? An Empirical Investigation of Twenty Software Startups

This study conducts a multiple case study on twenty European software startups and proposes a prototype-centric learning model in early stage software startups, and identifies factors that occur as barriers but also facilitators for prototyping in earlystage software startups.

Anh Nguyen-Duc, Xiaofeng Wang, P. Abrahamsson · 44 citations · ⚡5
#protein folding Open access Sep 2026

Programmable design of functional proteins from natural language

Pinal, a 16-billion-parameter foundation model that produces protein candidates from natural-language functional descriptions, supports natural language as a high-level interface for candidate generation in protein design, enabling programmable exploration with reduced reliance on manually specified structural or seque...

Fengyuan Dai, Shiyang You, Yudian Zhu et al. · 31 citations · ⚡3

Related blog posts

Google DeepMind Blog Sep 30, 2026

Introducing SynthID Bio

Proof of concept for watermarking AI-generated proteins while preserving biological function.

MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.