Jul 2026· Current Computer Science· Vol 05· 0 citations
TL;DR
The results of this study demonstrate that embeddings generated by protein language models contain disorder-relevant information and that classification methods based on similarities to other proteins can achieve similar performance levels as deep artificial neural networks, while also improving the ability to understand the meaning of the outputs and reducing the amount of computational resources needed to produce the desired results.
Abstract
Disordered proteins (IDPs) and disordered protein regions (IDRs) have
important roles in cellular signalling and regulation and in disease development. However, their
flexible conformations create challenges in annotating them through computational methods. Recent
advances in IDP disorder predictive methods based on deep learning models (e.g., SPOTDisorder,
AUCpreD, IDP-Fusion) have improved the accuracy of IDP disorder prediction; however,
all current methods still require heavy computational resources and are lacking in the ability
to interpret their results.
A new lightweight method for predicting IDRs called IDP-T5CKNN, which uses embeddings
produced by ProtT5-XL-UniRef50 and a cosine similarity k-nearest neighbour (KNN) classifier
to predict the residue-level disorder in an input protein sequence, is proposed in this study.
The residue level embeddings have been normalised and class-balanced, and evaluated using
standard binary classification metrics. This study also estimated the computational costs for the
IDP-T5CKNN method using formal Big-O notation.
For the independent MXD494 dataset, the IDP-T5CKNN method had a maximum correlation
coefficient (MCC) of 0.5459 and a balanced accuracy coefficient (BAC) of 0.7906, outperforming
all other currently available IDP disorder predictors. For the sample from the fiDPnn
Test176 dataset, the IDP-T5CKNN method produced an MCC of 0.3942 and a BAC of 0.7163, with
approximately equal sensitivity and specificity, while also achieving a comparable performance to
deep neural networks without requiring iterative training.
The results of this study demonstrate that embeddings generated by protein language
models contain disorder-relevant information and that classification methods based on similarities
to other proteins can achieve similar performance levels as deep artificial neural networks, while
also improving our ability to understand the meaning of the outputs and reducing the amount of
computational resources needed to produce the desired results.
Overall, the IDP-T5CKNN method provides a low-cost, scalable, and interpretable
method for making residue-level disorder predictions for entire proteomes.
Abstract Motivation Accurate identification of DNA-binding proteins (DBPs) and RNA-binding proteins (RBPs) is critical for elucidating transcriptional and post-transcriptional regulatory mechanisms. However, existing computational approaches often rely on inferred labels or domain-specific annotations, which limit the subsequent generalizability. Results This study aimed to introduce transformer-based classifiers for human DBPs and RBPs that rely solely on protein sequence information without engineered features or domain constraints. The models were implemented using ESM-2 with low-rank adaptation (LoRA) fine-tuning and trained on experimentally validated datasets, including chromatin immunoprecipitation sequencing (ChIP-seq) annotations for DBPs and eCLIP annotations for RBPs. Next, to evaluate biological relevance, we computed value-aware attention (VAT) scores aggregated across transformer layers to interpret model focus. In 20-fold cross-validation, the DBP model achieved an area under the receiver operating characteristic curve (AUROC) of 0.84 with a Matthews correlation coefficient (MCC) of 0.40, while the RBP model achieved an AUROC of 0.92 with an MCC of 0.46. Proteins predicted as nucleic acid-binding were enriched for known binding domains, and inspection of attention distributions revealed preferential focus on annotated functional regions rather than non-binding segments. These results demonstrate that attention-based protein language models can accurately identify nucleic acid-binding proteins directly from sequence data. Moreover, these models reveal biologically meaningful sequence determinants of binding, establishing an interpretable and scalable framework for proteome-wide characterization of protein–nucleic acid interactions. Availability and implementation Code is available on GitHub (https://github.com/CSB-hub/DRBP).
Hanjin Kim, Sung-Gwon Lee, Joo-Seong Oh et al.· Bioinformatics Advances· 0 citations
A neural network-based pipeline that integrates amino acid sequences with structural features is developed and provides a modular prototype for follow-up, more extensive protein modeling, including larger proteins and sequence of variable sizes.
Carl David Jasper Causin, M. Fyta· APL Machine Learning· 0 citations
The discordance persists: pLDDT correlates positively with PUNCH2 disorder in random and de novo proteins and negatively with β-strand fraction, opposite to the conserved and disordered baselines, a concrete failure mode that protein designers and other working on sequences remote in sequence space should be aware of when relying on predictor outputs.
Lars A. Eicholt, Lasse Middendorf· bioRxiv· 0 citations
The fundamental cellular processes, including transcriptional regulation, chromatin organization, and genome maintenance, are regulated by DNA-binding proteins (DBPs). Mutations in DBPs can alter protein-DNA interactions, leading to tumor development. However, identifying such driver mutations remains a major challenge due to limitations of experimental approaches.
We have trained a machine learning model, DBP-CanPred, to identify driver mutations in DBPs. We used the sequence-derived evolutionary features, as well as structure-based features such as mutation-perturbed structural descriptors.
We evaluated DBP-CanPred using a curated test set, achieving an AU-ROC of 0.86 and a balanced accuracy of 0.79. Further analysis based on substitution-type showed consistent performance across different categories, especially higher performance on charged residues. In addition, we applied the model on an independent dataset and identified potential driver mutations with high confidence scores.
The study contributes to understanding mutation patterns in DNA-binding proteins and supports variant interpretation in cancer research.
A. Phogat, Sowmya Ramaswamy Krishnan, Medha Pandey et al.· Frontiers in Bioinformatics· 0 citations
DHST is proposed, a deep hybrid structure–topology framework that integrates sequence semantics from a pretrained protein language model with local structural information learned by a residual graph convolutional network and introduces site-specific persistent homology to encode multi-scale topological invariants and a topology-guided residue-wise gated fusion module to modulate structure–semantics representations using local topological embeddings.
Bin Lu, Fujun Xiang, Hai-Long Wang et al.· Applied Sciences· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.