Five machine learning classifiers are benchmarked on a real, published GUIDE-seq off-target dataset and the low absolute precision achievable in this severely imbalanced, small-positive-class setting is reported, as a realistic picture of what off-target classifiers can and cannot yet deliver from sequence alone.
Abstract
Off-target cleavage is a central safety concern for CRISPR-Cas9 genome editing, particularly in therapeutic applications where unintended double-strand breaks carry clinical risk. We benchmarked five machine learning classifiers — logistic regression on mismatch-count summary features, a random forest and a gradient boosting model on one-hot-encoded sgRNA/candidate-site sequence pairs, a one-dimensional convolutional neural network (CNN) over the positional mismatch map, and a gradient-boosting/CNN ensemble — on a real, published GUIDE-seq off-target dataset (Kleinstiver et al., 2016, Nature) comprising 95,829 candidate off-target sites for five sgRNAs, of which only 54 (0.06%) were experimentally validated as true cleavage sites. On a held-out, stratified test split (n = 19,166; 11 true positives), gradient boosting on combined mismatch and sequence features performed best (ROC-AUC = 0.997, PR-AUC = 0.355, best F1 = 0.50), outperforming a random forest on raw sequence encoding alone (PR-AUC = 0.083) and a sequence CNN (PR-AUC = 0.129). Because the positive class is extremely rare, we report precision-recall AUC as the primary metric rather than ROC-AUC, which is inflated by the large negative class. A positional mismatch analysis showed that experimentally validated off-target sites carried substantially fewer mismatches overall than non-cleaved candidate sites (mean 3.6 vs. 5.9 mismatches across the 23-nucleotide target), and were markedly more mismatch-intolerant in the 10-nucleotide PAM-proximal seed region (11.3% vs. 27.4% per-position mismatch rate) and at the PAM itself (6.8% vs. 16.0%), consistent with established seed-region and PAM-sensitivity models of Cas9 target recognition. We report these findings, including the low absolute precision achievable in this severely imbalanced, small-positive-class setting, as a realistic picture of what off-target classifiers can and cannot yet deliver from sequence alone.
Abstract Summary CRISPR-Cas9 has become a widely used tool for genome editing. However, its off-target cleavage caused by partial sequence matches with guide RNAs (gRNAs) remains a critical limitation. Recently, abasic gRNAs (ØXØ) have been developed to enhance target specificity, but their effects vary depending on the positional sequence context. Here, we present abCRISPR, a deep neural network (DNN) framework for the rational design of ØXØ sequences with minimized off-target activity. abCRISPR leverages informative few-shot training with paired datasets of abasic and unmodified gRNAs, using high-quality random mismatch target libraries, exhaustively sequenced for mismatched off-target substrates (n = 97583) in in vitro CRISPR-Cas9 cleavage experiments. Predicted off-target activities for both abasic and unmodified gRNAs showed strong correlation with experimental data (r ≥ 0.95, 10-fold cross-validation). Notably, these comprehensive training sets provide robust ground-truth negatives, enabling accurate and sensitive prediction of off-targets. For unmodified gRNAs, abCRISPR (AUC = 0.98) was validated to outperform existing deep learning-based methods (AUC = 0.45–0.68). When applied to the human genome, abCRISPR generated ØXØ sequences, covering 58 875 004 potent CRISPR-targetable sites with improved target specificity. Together, this work provides a comprehensive bioinformatics resource for safe and precise CRISPR-Cas9 genome editing. Availability and implementation The source code for abCRISPR and training data are available at https://doi.org/10.5281/zenodo.20398246. abCRISPR results for the human genome are available at http://clip.korea.ac.kr/abCRISPR/
Geun-Woo D. Kim, Dowoon Gu, Mingyo Park et al.· Bioinform.· 0 citations
Accurate identification of CRISPR-Cas9 off-target sites is essential for the safety assessment of genome-editing-based therapies. While numerous in silico prediction tools have been developed, their comparative performance and practical utility in preclinical workflows remain incompletely defined. We performed a systematic benchmarking of 14 in silico CRISPR-Cas9 off-target prediction tools, including both standard approaches and machine learning-based models. The analysis was based on a curated dataset derived from the CRISPRoffT database, comprising 3,827 deep-sequenced genomic sites across 26 guide RNA/Cas9 combinations in human cells. Sites with indel frequencies ≥0.1% were operationally defined as true off-targets. We evaluated tool performance using score distributions, correlation with indel frequencies, precision-recall characteristics, recall among top-ranked candidate sites, and the effect of combining tools. All tools assigned higher scores to true off-target sites compared with nontarget sites, although substantial overlap between classes was observed. Correlation between prediction scores and indel frequencies was weak to moderate, indicating limited ability to predict editing magnitude. Precision-recall performance was moderate across all tools, reflecting inherent trade-offs between sensitivity and specificity. Recall increased with the number of predicted sites considered, reaching approximately 77% among the top 500 and up to 83% among the top 1,250 sites, but leaving a substantial fraction of true off-targets undetected. Combining tools yielded only modest improvements. Current in silico tools enable prioritization of CRISPR-Cas9 off-target candidates but remain limited in their ability to comprehensively identify and quantitatively predict off-target activity. Our findings highlight the importance of considering both ranking performance and candidate site coverage and support the use of combined computational and experimental strategies for robust off-target assessment in preclinical gene editing workflows.
M. M. Kaufmann, Maren Hackenberg, William Jobson Pargeter et al.· Human Gene Therapy· 0 citations
Due to their high efficiency and programmability, CRISPR/Cas systems are commonly used in gene function experiments and gene therapy, but off-target effects are still a big problem-Cas proteins may cut non-target sites. Most machine learning and deep learning methods for predicting the off-target activity of gRNA–DNA pairs are point predictions, which are hard to quantify and weaken the validity of the assessment and amplify the risk of decision-making in high-risk scenarios. Therefore, uncertainty quantification approaches have been increasingly applied into the research of this field. This paper reviews the related works on applying uncertainty modeling in this field. Specifically, the concepts of data uncertainty and model uncertainty and their modeling techniques will be reviewed, including heteroscedastic regression, probabilistic predictive models, Deep Ensembles and MC Dropout. Furthermore, this paper will introduce some key metrics for the evaluation of the reliability of uncertainty models, including predictive accuracy, confidence interval calibration and performance on out-of-distribution (OOD) data. By reviewing the related studies, this paper concludes that it is very important to use uncertainty-aware models to improve the reliability of CRISPR predictions in this field.
Track-seq2 provides a robust platform for sensitive off-target detection in primary cells, with sensitivity comparable to or exceeding current state-of-the-art methods.
ABSTRACT Amplification‐free Cas12a diagnostics with split crRNA enable rapid and programmable target recognition, yet insufficient understanding of DNA activator architecture prevents predictable control over trans‐cleavage activity and sensitivity. Here we systematically map over 200 split DNA activator configurations by introducing nicks at every position across both strands. The mapping reveals that target strand nicks suppress activity with position‐dependent severity, while non‐target strand nicks enhance activity. Guided by these rules, we engineer an optimized split activator pair that achieves attomolar microRNA detection (LOD: 112 aM), ∼480‐fold higher sensitivity than intact activators. The enhanced sensitivity supports multiplexed live‐cell profiling of five miRNAs for machine learning‐based cancer cell stratification, and is further generalized to non‐nucleic acid targets, including APE1 enzyme (0.0073 U/L) and HClO (2.37 pM) through position‐informed cleavable sites. This work provides a generalizable methodology for engineering CRISPR‐Cas12a performance across diagnostic and biosensing applications.
Xiaoyan Tang, Zhe Li, Yuning Lu et al.· Advancement of science· 0 citations
Traditional CRISPR-Cas12a mutation detection systems are limited by poor single-base specificity, target-specific crRNA redesign, and insufficient sensitivity for low-abundance mutations, restricting their clinical liquid biopsy applications. Herein, we developed a crRNA-universal, sensitive and specific CRISPR-Cas12a detection platform, termed DESIC (double-end blocker and split-input mediated CRISPR-Cas12a system), for single-base mutation detection. The DESIC system adopts two key structural designs: double-end blocker (DEB) and duplicated split-input (SIN). The DEB spatially isolates crRNA recognition and target-binding regions, enabling universal detection of various mutation sites without crRNA redesign. The SIN strategy amplifies thermodynamic differences from single-base mismatches, greatly improving single-nucleotide discrimination. We targeted four prevalent pancreatic cancer KRAS mutations (G12D, G12R, G12V, Q61H) and optimized the system to achieve optimal discrimination. The optimized DESIC system exhibited ultra-low limits of detection down to 0.01% mutant allele fraction with reliable linear quantitative performance. Clinical validation using 15 pairs of pancreatic cancer tissue and peripheral blood samples confirmed that DESIC results were highly consistent with gold-standard NGS data. With a flexible modular design, this low-cost, easy-operated platform can be readily extended to multiple tumor mutations, holding great potential for tumor liquid biopsy and early molecular diagnosis.
Shi-Zhen Li, Yangwei Liao, Xiaoxiang Wang et al.· Biosensors & bioelectronics· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.