Skip to content
Open access

DNAreader: accurate prediction of DNA-binding residues in structured and disordered proteins using transformers and contrastive learning

Aug 2026 · Nucleic Acids Research · Vol 54 · 0 citations · 103 references
Medicine

TL;DR

This work introduces DNAreader, the first predictor specifically designed to predict DBRs in the structured and disordered sequence regions, and develops the DNAreaderDBIDR module, which accurately predicts DNA-binding IDRs, providing flexibility to identify DBRs within IDRs or to predict entire disordered DNA-binding regions.

Abstract

Abstract Accurate predictions of DNA-binding residues (DBRs) in protein sequences facilitate decoding molecular-level mechanisms underlying cellular functions that involve protein–DNA interactions. While dozens of these predictors have been released, they target either structured or intrinsically disordered regions (IDRs), and the latter were trained to predict less detailed DNA-binding IDRs rather than DBRs. Given this dichotomy, the structure-trained methods underperform on disordered proteins, and vice versa. Moreover, they suffer from high cross-prediction rates, incorrectly labeling many residues that interact with non-DNA ligands as DBRs. We address these issues by introducing DNAreader, the first predictor specifically designed to predict DBRs in the structured and disordered sequence regions. DNAreader relies on an innovative stacked transformer encoder network that combines batch training and contrastive learning, which substantially boosts predictive performance. Using two low-similarity test datasets, we demonstrate that DNAreader statistically outperforms existing tools, performs well for structured and disordered regions, and produces very few cross-predictions. We also developed the DNAreaderDBIDR module, which accurately predicts DNA-binding IDRs, providing flexibility to identify DBRs within IDRs or to predict entire disordered DNA-binding regions. We release DNAreader as a user-friendly web server at http://biomine.cs.vcu.edu/servers/DNAreader/, with the corresponding source code at https://github.com/jianzhang-xynu/DNAreader.

Read PDF

Similar papers

Open access Aug 2026

DBP-CanPred: a machine learning model for predicting cancer-causing mutations in DNA-binding proteins

The fundamental cellular processes, including transcriptional regulation, chromatin organization, and genome maintenance, are regulated by DNA-binding proteins (DBPs). Mutations in DBPs can alter protein-DNA interactions, leading to tumor development. However, identifying such driver mutations remains a major challenge due to limitations of experimental approaches. We have trained a machine learning model, DBP-CanPred, to identify driver mutations in DBPs. We used the sequence-derived evolutionary features, as well as structure-based features such as mutation-perturbed structural descriptors. We evaluated DBP-CanPred using a curated test set, achieving an AU-ROC of 0.86 and a balanced accuracy of 0.79. Further analysis based on substitution-type showed consistent performance across different categories, especially higher performance on charged residues. In addition, we applied the model on an independent dataset and identified potential driver mutations with high confidence scores. The study contributes to understanding mutation patterns in DNA-binding proteins and supports variant interpretation in cancer research.

A. Phogat, Sowmya Ramaswamy Krishnan, Medha Pandey et al. · 0 citations
Open access Aug 2026

A discrete protein subset drives structure prediction discordance in orphan proteins

The discordance persists: pLDDT correlates positively with PUNCH2 disorder in random and de novo proteins and negatively with β-strand fraction, opposite to the conserved and disordered baselines, a concrete failure mode that protein designers and other working on sequences remote in sequence space should be aware of when relying on predictor outputs.

Lars A. Eicholt, Lasse Middendorf · 0 citations
Open access Jul 2026

Predictions of protein–protein interactions: Learning sequences and structures

A neural network-based pipeline that integrates amino acid sequences with structural features is developed and provides a modular prototype for follow-up, more extensive protein modeling, including larger proteins and sequence of variable sizes.

Carl David Jasper Causin, M. Fyta · 0 citations
Open access Aug 2026

DNCLA: A Deep Learning Model for TFBS Identification Based on Structural and Conformational Properties of Nucleotides and Dinucleotides

Identifying transcription factor binding sites (TFBSs) is fundamental to understanding complex gene regulatory mechanisms and the functions of non-coding regions. Although existing methods have achieved substantial strides, capturing both local structural features and long-range spatial dependencies within DNA sequences remains a major challenge for improving prediction accuracy. In this study, we propose DNCLA, a deep learning model that synergizes multisize convolutional fusion, Bidirectional Long ShortTerm Memory (Bi-LSTM) networks, and a multi-head self-attention mechanism. At the feature extraction level, DNCLA breaks through the limitations of traditional single-sequence encoding by fusing Nucleotide Chemical Properties (NCP) with Dinucleotide Physicochemical Properties (DPCP). NCP provides a refined characterization of chemical differences between bases based on ring structures, hydrogen bond sites, and functional group properties, while DPCP introduces parameters such as local structural stability and geometric flexibility of the DNA. Subsequently, the model extracts spatial evolution from these high-dimensional features through a multi-size convolutional module; captures long-range spatial dependencies using Bi-LSTM layers; and employs a multi-head self-attention mechanism to achieve adaptive weight distribution of global features, thereby enhancing the perception of key regulatory motifs. Results from training and testing the proposed model on 165 ChIPseq datasets demonstrate that DNCLA possesses robust generalization capabilities and high predictive performance in TFBSs identification. This suggests that the incorporation of physicochemical features better elucidates the essence of interactions between transcription factors and DNA.

Jingjue Wei, Jie Feng · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.