Skip to content

IDP-T5CKNN: Protein Language Models for Prediction of Intrinsically Disordered Proteins

Jul 2026 · Current Computer Science · Vol 05 · 0 citations

TL;DR

The results of this study demonstrate that embeddings generated by protein language models contain disorder-relevant information and that classification methods based on similarities to other proteins can achieve similar performance levels as deep artificial neural networks, while also improving the ability to understand the meaning of the outputs and reducing the amount of computational resources needed to produce the desired results.

Abstract

Disordered proteins (IDPs) and disordered protein regions (IDRs) have important roles in cellular signalling and regulation and in disease development. However, their flexible conformations create challenges in annotating them through computational methods. Recent advances in IDP disorder predictive methods based on deep learning models (e.g., SPOTDisorder, AUCpreD, IDP-Fusion) have improved the accuracy of IDP disorder prediction; however, all current methods still require heavy computational resources and are lacking in the ability to interpret their results. A new lightweight method for predicting IDRs called IDP-T5CKNN, which uses embeddings produced by ProtT5-XL-UniRef50 and a cosine similarity k-nearest neighbour (KNN) classifier to predict the residue-level disorder in an input protein sequence, is proposed in this study. The residue level embeddings have been normalised and class-balanced, and evaluated using standard binary classification metrics. This study also estimated the computational costs for the IDP-T5CKNN method using formal Big-O notation. For the independent MXD494 dataset, the IDP-T5CKNN method had a maximum correlation coefficient (MCC) of 0.5459 and a balanced accuracy coefficient (BAC) of 0.7906, outperforming all other currently available IDP disorder predictors. For the sample from the fiDPnn Test176 dataset, the IDP-T5CKNN method produced an MCC of 0.3942 and a BAC of 0.7163, with approximately equal sensitivity and specificity, while also achieving a comparable performance to deep neural networks without requiring iterative training. The results of this study demonstrate that embeddings generated by protein language models contain disorder-relevant information and that classification methods based on similarities to other proteins can achieve similar performance levels as deep artificial neural networks, while also improving our ability to understand the meaning of the outputs and reducing the amount of computational resources needed to produce the desired results. Overall, the IDP-T5CKNN method provides a low-cost, scalable, and interpretable method for making residue-level disorder predictions for entire proteomes.

View source

Similar papers

Open access Aug 2026

Interpretable prediction of nucleic acid-binding proteins using a protein language model

Abstract Motivation Accurate identification of DNA-binding proteins (DBPs) and RNA-binding proteins (RBPs) is critical for elucidating transcriptional and post-transcriptional regulatory mechanisms. However, existing computational approaches often rely on inferred labels or domain-specific annotations, which limit the subsequent generalizability. Results This study aimed to introduce transformer-based classifiers for human DBPs and RBPs that rely solely on protein sequence information without engineered features or domain constraints. The models were implemented using ESM-2 with low-rank adaptation (LoRA) fine-tuning and trained on experimentally validated datasets, including chromatin immunoprecipitation sequencing (ChIP-seq) annotations for DBPs and eCLIP annotations for RBPs. Next, to evaluate biological relevance, we computed value-aware attention (VAT) scores aggregated across transformer layers to interpret model focus. In 20-fold cross-validation, the DBP model achieved an area under the receiver operating characteristic curve (AUROC) of 0.84 with a Matthews correlation coefficient (MCC) of 0.40, while the RBP model achieved an AUROC of 0.92 with an MCC of 0.46. Proteins predicted as nucleic acid-binding were enriched for known binding domains, and inspection of attention distributions revealed preferential focus on annotated functional regions rather than non-binding segments. These results demonstrate that attention-based protein language models can accurately identify nucleic acid-binding proteins directly from sequence data. Moreover, these models reveal biologically meaningful sequence determinants of binding, establishing an interpretable and scalable framework for proteome-wide characterization of protein–nucleic acid interactions. Availability and implementation Code is available on GitHub (https://github.com/CSB-hub/DRBP).

Hanjin Kim, Sung-Gwon Lee, Joo-Seong Oh et al. · 0 citations
Open access Jul 2026

Predictions of protein–protein interactions: Learning sequences and structures

A neural network-based pipeline that integrates amino acid sequences with structural features is developed and provides a modular prototype for follow-up, more extensive protein modeling, including larger proteins and sequence of variable sizes.

Carl David Jasper Causin, M. Fyta · 0 citations
Open access Aug 2026

A discrete protein subset drives structure prediction discordance in orphan proteins

The discordance persists: pLDDT correlates positively with PUNCH2 disorder in random and de novo proteins and negatively with β-strand fraction, opposite to the conserved and disordered baselines, a concrete failure mode that protein designers and other working on sequences remote in sequence space should be aware of when relying on predictor outputs.

Lars A. Eicholt, Lasse Middendorf · 0 citations
Open access Aug 2026

DBP-CanPred: a machine learning model for predicting cancer-causing mutations in DNA-binding proteins

The fundamental cellular processes, including transcriptional regulation, chromatin organization, and genome maintenance, are regulated by DNA-binding proteins (DBPs). Mutations in DBPs can alter protein-DNA interactions, leading to tumor development. However, identifying such driver mutations remains a major challenge due to limitations of experimental approaches. We have trained a machine learning model, DBP-CanPred, to identify driver mutations in DBPs. We used the sequence-derived evolutionary features, as well as structure-based features such as mutation-perturbed structural descriptors. We evaluated DBP-CanPred using a curated test set, achieving an AU-ROC of 0.86 and a balanced accuracy of 0.79. Further analysis based on substitution-type showed consistent performance across different categories, especially higher performance on charged residues. In addition, we applied the model on an independent dataset and identified potential driver mutations with high confidence scores. The study contributes to understanding mutation patterns in DNA-binding proteins and supports variant interpretation in cancer research.

A. Phogat, Sowmya Ramaswamy Krishnan, Medha Pandey et al. · 0 citations
Open access Aug 2026

DHST: A Deep Hybrid Structure–Topology Framework for Accurate Protein Function Prediction

DHST is proposed, a deep hybrid structure–topology framework that integrates sequence semantics from a pretrained protein language model with local structural information learned by a residual graph convolutional network and introduces site-specific persistent homology to encode multi-scale topological invariants and a topology-guided residue-wise gated fusion module to modulate structure–semantics representations using local topological embeddings.

Bin Lu, Fujun Xiang, Hai-Long Wang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.