Symphony-Bind: Prediction of Protein Binding Sites for 11 Representative Small Molecules and Ions via Fine-Tuning Protein Language Models and Grouped Multi-Task Learning
Aug 2026· ACS Omega· Vol 11, pp. 49447 - 49463· 0 citations· 72 references
Medicine
TL;DR
A Grouped Multi-Task Learning (GMTL) strategy is implemented, allowing the model to capture shared binding patterns among ligands with similar biological significance, allowing the model to capture shared binding patterns among ligands with similar biological significance.
Abstract
Accurately identifying protein binding sites for small molecules and ions is crucial for understanding biological processes and advancing drug discovery. Pretrained protein language models (pLMs) have emerged as powerful tools for this purpose, but existing prediction models often face a trade-off when using pLMs: freezing pLMs limits their adaptability, while fully fine-tuning them requires high computational costs. To address this trade-off, we attempted to introduce Parameter-Efficient Fine-Tuning (PEFT) as a promising solution to balance efficiency and performance. Specifically, we applied five PEFT strategies (i.e., LoRA, QLoRA, DoRA, AdaLoRA, and IA3) to fine-tune four pLMs (i.e., ProtT5, ProtBERT, ESM2–150M, and ESM2–650M) and assessed their predictive performance across protein binding site tasks for 11 representative small molecules and ions. The results clearly indicate that the LoRA-enhanced ESM2–650M consistently outperforms all other combinations. Despite this robust baseline, training independent models for specific small molecules remains challenging due to the scarcity of high-quality binding data. To bridge this gap, we implemented a Grouped Multi-Task Learning (GMTL) strategy, allowing the model to capture shared binding patterns among ligands with similar biological significance. Experimental results demonstrate that this strategy significantly enhances predictive performance. Building upon these insights, we present Symphony-Bind. It is a GMTL framework that leverages LoRA-enhanced ESM2–650M to extract embeddings, which are subsequently refined by a shared ConvBERT module and then processed by ligand-specific MLPs for precise binding site prediction. Performance evaluation on 11 representative ligand tasks shows that Symphony-Bind achieves average MCC values of 0.561, 0.629, and 0.324 for the nucleotide, cofactor, and inorganic ion groups, surpassing evaluated sequence-based state-of-the-art methods while remaining competitive with structure-based models.
Testing the ability of common large language models to consider design principles to generate de novo proteins that bind metals and lipophilic small molecules without copying existing sequences highlights the utility of LLMs in making protein design more comprehensible and accessible to users without sophisticated design expertise.
Protein language models (PLMs) provide powerful representations of protein sequence, but their utility for proteome-scale binding-site retrieval remains unclear. Here, we present PocketScope, a training-free framework that represents cavity-lining residues using frozen ESM-C 600M embeddings and retrieves related binding sites through exhaustive lateinteraction MaxSim, without pooling or approximate nearest-neighbor search. PocketScope identified 153,805 cavities across 37,682 proteins in the AlphaFold human proteome and recovered documented drug off-targets across a curated set of pharmacological pairs. On the ProSPECCTs benchmark, PocketScope ranks 1st of 23 methods by mean rank across the ten collections. PocketScope provides a practical framework for proteome-scale off-target prediction. PocketScope is open source and also freely available as a web server at https://www.bhargavaresearch.org/pocketscope.
Post-translational modifications (PTMs) regulate protein function, making accurate residue-level PTM prediction essential for understanding cellular mechanisms and disease pathways. While decoder-only protein language models (PLMs) pretrained with the causal language modeling (CLM) objective have driven breakthroughs across various bioinformatics tasks, their potential for PTM prediction remains largely underexplored. CLM-based PLMs that rely on Byte-Pair Encoding (BPE) for tokenization, such as ProtGPT2, introduce intra-token label collision by merging multiple amino acids with conflicting labels into a single token, creating a major bottleneck for residue-level tasks. To overcome this, we propose TaHL-PTM (Target-Hooked Low-rank adaptation for PTM prediction), a novel framework that integrates target-hooked tokenization with site-directed discriminative LoRA fine-tuning. Target-hooked tokenization constrains tokenization around the candidate residue using dedicated marker tokens to eliminate intra-token label collision while preserving the surrounding sequence context, whereas the proposed discriminative objective repurposes the standard generative CLM objective for residue-level PTM classification by directly optimizing the separation between modified and unmodified sites. We benchmark TaHL-PTM across six distinct PTM tasks on ProtGPT2 and ProGen2 models. TaHL-PTM consistently improves MCC, with the largest gain of up to +0.11 for tyrosine phosphorylation (0.34 to 0.45), alongside improvements in F1, AUROC, and AUPR. Performance gains are more pronounced for collision-affected samples, validating the effectiveness of target-hooked tokenization, while consistent improvements across both BPE-based and per-residue-based causal PLMs demonstrate that the proposed framework generalizes across models with different pretraining tokenization schemes.
Bhawana Prasain, Pawel Pratyush, Stefan Schulze et al.· bioRxiv· 0 citations
MBPBERT provides a scalable and efficient in silico solution for high-throughput discovery of novel MBPs and screening of peptides with metal-specific binding preferences, potentially reducing the reliance on resource-intensive experimental validation.
“NextTopDocker” is presented, a large, up-to-date, open-access data set for docking-power assessment comprising 14,038 training and 5201 test entries across 3173 unique protein targets, constructed from the Protein Data Bank.
Cao-Minh Truong, Pedro J. Ballester, O. Taboureau et al.· Journal of Medicinal Chemist...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.