Skip to content
Open access

Symphony-Bind: Prediction of Protein Binding Sites for 11 Representative Small Molecules and Ions via Fine-Tuning Protein Language Models and Grouped Multi-Task Learning

Aug 2026 · ACS Omega · Vol 11, pp. 49447 - 49463 · 0 citations · 72 references
Medicine

TL;DR

A Grouped Multi-Task Learning (GMTL) strategy is implemented, allowing the model to capture shared binding patterns among ligands with similar biological significance, allowing the model to capture shared binding patterns among ligands with similar biological significance.

Abstract

Accurately identifying protein binding sites for small molecules and ions is crucial for understanding biological processes and advancing drug discovery. Pretrained protein language models (pLMs) have emerged as powerful tools for this purpose, but existing prediction models often face a trade-off when using pLMs: freezing pLMs limits their adaptability, while fully fine-tuning them requires high computational costs. To address this trade-off, we attempted to introduce Parameter-Efficient Fine-Tuning (PEFT) as a promising solution to balance efficiency and performance. Specifically, we applied five PEFT strategies (i.e., LoRA, QLoRA, DoRA, AdaLoRA, and IA3) to fine-tune four pLMs (i.e., ProtT5, ProtBERT, ESM2–150M, and ESM2–650M) and assessed their predictive performance across protein binding site tasks for 11 representative small molecules and ions. The results clearly indicate that the LoRA-enhanced ESM2–650M consistently outperforms all other combinations. Despite this robust baseline, training independent models for specific small molecules remains challenging due to the scarcity of high-quality binding data. To bridge this gap, we implemented a Grouped Multi-Task Learning (GMTL) strategy, allowing the model to capture shared binding patterns among ligands with similar biological significance. Experimental results demonstrate that this strategy significantly enhances predictive performance. Building upon these insights, we present Symphony-Bind. It is a GMTL framework that leverages LoRA-enhanced ESM2–650M to extract embeddings, which are subsequently refined by a shared ConvBERT module and then processed by ligand-specific MLPs for precise binding site prediction. Performance evaluation on 11 representative ligand tasks shows that Symphony-Bind achieves average MCC values of 0.561, 0.629, and 0.324 for the nucleotide, cofactor, and inorganic ion groups, surpassing evaluated sequence-based state-of-the-art methods while remaining competitive with structure-based models.

Read PDF

Similar papers

#protein folding Open access Sep 2026

De novo design of ligand binding proteins using large language models alone

Testing the ability of common large language models to consider design principles to generate de novo proteins that bind metals and lipophilic small molecules without copying existing sequences highlights the utility of LLMs in making protein design more comprehensible and accessible to users without sophisticated design expertise.

Unknown authors · 0 citations
Open access Sep 2026

Retrieval of binding sites across the AlphaFold human proteome using protein language model representations

Protein language models (PLMs) provide powerful representations of protein sequence, but their utility for proteome-scale binding-site retrieval remains unclear. Here, we present PocketScope, a training-free framework that represents cavity-lining residues using frozen ESM-C 600M embeddings and retrieves related binding sites through exhaustive lateinteraction MaxSim, without pooling or approximate nearest-neighbor search. PocketScope identified 153,805 cavities across 37,682 proteins in the AlphaFold human proteome and recovered documented drug off-targets across a curated set of pharmacological pairs. On the ProSPECCTs benchmark, PocketScope ranks 1st of 23 methods by mean rank across the ten collections. PocketScope provides a practical framework for proteome-scale off-target prediction. PocketScope is open source and also freely available as a web server at https://www.bhargavaresearch.org/pocketscope.

Unknown authors · 0 citations
Open access Aug 2026

TaHL-PTM: Post-Translational Modification Prediction in Proteins via Target-Hooked Discriminative Fine-Tuning of Decoder-only Protein Language Models

Post-translational modifications (PTMs) regulate protein function, making accurate residue-level PTM prediction essential for understanding cellular mechanisms and disease pathways. While decoder-only protein language models (PLMs) pretrained with the causal language modeling (CLM) objective have driven breakthroughs across various bioinformatics tasks, their potential for PTM prediction remains largely underexplored. CLM-based PLMs that rely on Byte-Pair Encoding (BPE) for tokenization, such as ProtGPT2, introduce intra-token label collision by merging multiple amino acids with conflicting labels into a single token, creating a major bottleneck for residue-level tasks. To overcome this, we propose TaHL-PTM (Target-Hooked Low-rank adaptation for PTM prediction), a novel framework that integrates target-hooked tokenization with site-directed discriminative LoRA fine-tuning. Target-hooked tokenization constrains tokenization around the candidate residue using dedicated marker tokens to eliminate intra-token label collision while preserving the surrounding sequence context, whereas the proposed discriminative objective repurposes the standard generative CLM objective for residue-level PTM classification by directly optimizing the separation between modified and unmodified sites. We benchmark TaHL-PTM across six distinct PTM tasks on ProtGPT2 and ProGen2 models. TaHL-PTM consistently improves MCC, with the largest gain of up to +0.11 for tyrosine phosphorylation (0.34 to 0.45), alongside improvements in F1, AUROC, and AUPR. Performance gains are more pronounced for collision-affected samples, validating the effectiveness of target-hooked tokenization, while consistent improvements across both BPE-based and per-residue-based causal PLMs demonstrate that the proposed framework generalizes across models with different pretraining tokenization schemes.

Bhawana Prasain, Pawel Pratyush, Stefan Schulze et al. · 0 citations
Aug 2026

MBPBERT: A Large Language Model for Metal-Binding Peptide Discovery.

MBPBERT provides a scalable and efficient in silico solution for high-throughput discovery of novel MBPs and screening of peptides with metal-specific binding preferences, potentially reducing the reliance on resource-intensive experimental validation.

Guifen Jian, Xin-Wei Li, Yu Chen et al. · 0 citations
Open access Aug 2026

NextTopDocker: A Large-Scale Docking-Power Benchmark Reveals Limitations of Current End-to-End Machine-Learning Docking and the Strength of Hybrid Rescoring.

“NextTopDocker” is presented, a large, up-to-date, open-access data set for docking-power assessment comprising 14,038 training and 5201 test entries across 3173 unique protein targets, constructed from the Protein Data Bank.

Cao-Minh Truong, Pedro J. Ballester, O. Taboureau et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.