Skip to content

MT-ProtBERT: Multi-task Learning ProtBERT for Intrinsically Disordered Proteins Classification with Scarce Data

Sep 2026 · 0 citations · 84 references
Computer Science

TL;DR

MT-ProtBERT consistently outperforms PARROT, an RNN-based IDP-specific model, across all tasks, and demonstrates that combining self-supervised and biochemistry-informed tasks, and multi-scale learning enables robust modeling of unstructured proteins under data scarcity.

Abstract

Intrinsically disordered proteins (IDPs) differ from folded proteins in that they are dynamic, lack a stable three-dimensional conformation, and have low sequence similarity between similar proteins. The conformational heterogeneity of IDPs - while beneficial for their diverse functions - limits the use of traditional experimental tools to determine their conformation. The experimental difficulty, along with low sequence similarity, results in data scarcity, and makes it difficult to classify/detect IDPs that are similar or dissimilar, a task relevant to understand biology and evolution. We address this challenge using Multi-task ProtBERT (MT-ProtBERT), a multi-task extension of ProtBERT tailored for low-data regimes. MT-ProtBERT integrates Dynamic Window Masking, a Multi-Scale 1D Convolutional classifier (MS-Conv1D), and auxiliary objectives that jointly optimize masked language modeling and biochemistry-informed tasks. We evaluate this framework on two tasks under limited data: (i) phosphorylation site prediction (S/T/Y) in short sequences and small datasets, and (ii) protein compaction prediction on two small datasets (684 and 530 sequences), including sequences comparable in length to typical disordered regions. MT-ProtBERT consistently outperforms PARROT, an RNN-based IDP-specific model, across all tasks. These results demonstrate that combining self-supervised and biochemistry-informed tasks, and multi-scale learning enables robust modeling of unstructured proteins under data scarcity.

View source

Similar papers

Open access Aug 2026

FuncSeek: Multi-PLM contrastive learning for protein functional similarity search

FuncSeek is described, a contrastive learning model which utilizes three diverse, complementary PLMs: ESM2 (to model evolutionary co-variation), ProstT5 (for bilingual sequence and structure embeddings), and ProteinBERT (for functional semantic similarities) that each capture a different aspect of protein biology: evol...

Leendert J. Cloete, Hugh G. Patterton · 0 citations
Open access Aug 2026

3Dloop-FPSSM: Predicting Protein–Protein Interactions by Fusing 3D Local Optimal Oriented Pattern and Folded Position-Specific Scoring Matrix

This study introduces a novel sequence-based framework for PPI prediction, which combines position-specific scoring matrices (PSSMs), 3D local optimal orientation patterns (3Dloop), and histogram gradient boosting (HistGB) and shows that the approach provides a reliable and efficient solution for PPI prediction.

Fangfang Bai, Dan Liu, Guangxian Wang et al. · 0 citations
Open access Aug 2026

A discrete protein subset drives structure prediction discordance in orphan proteins

The discordance persists: pLDDT correlates positively with PUNCH2 disorder in random and de novo proteins and negatively with β-strand fraction, opposite to the conserved and disordered baselines, a concrete failure mode that protein designers and other working on sequences remote in sequence space should be aware of w...

Lars A. Eicholt, Lasse Middendorf · 0 citations
Open access Sep 2026

NeoToxPred: a fine-tuned protein language model with orthologous and length-stratified negative controls for robust toxicity classification

Although protein toxins represent valuable pharmacological templates, predicting toxicity directly from primary sequences is inherently challenging because of their evolutionary dynamics. Active toxins and their benign homologues frequently share identical structural scaffolds and differ by only a few key residue subst...

Seongmin Kim, Min-Seok Kim, Chungoo Park · 0 citations
Open access Sep 2026

Supervised Protein Structure Classification Using Topological Persistence With DeltaFold.

The DeltaFold Classifier (DFC) is introduced, a fast, alignment-free, protein structure classification pipeline based on topological data analysis that achieves performance comparable to that of structure-based comparison methods while substantially improving computational efficiency.

Joseph Nardin-Gennequin, Gabriela Ciuperca, Céline Brochier-Armanet et al. · 0 citations

Related blog posts

GPT-Lab Sep 3, 2026

Adaptive AI Agents in Construction Workflows

Adaptive AI agents can help make BIM data more machine-readable by navigating IFC models, interpreting inconsistent information, and mapping it to defined standards. In this blog, Alok Rawat shares findings from a real-world pilot in construction workflows. The post Adaptive AI Agents in Construction Workflows appeared first on GPT-Lab.

GPT-Lab Aug 28, 2026

We built an AI factory for HVAC control

What does it take to trust AI-driven HVAC optimization? Our AI Model Factory combines agents, machine learning, reinforcement learning and deterministic checks in a governed workflow designed for messy, real-world building data. The post We built an AI factory for HVAC control appeared first on GPT-Lab.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.