Skip to content

GPCR-SLM: Small Language Model-Based Classification of GPCRs Using Knowledge Distillation Technique.

Jul 2026 · IEEE transactions on computational biology and bioinformatics · Vol PP, pp. 1-14 · 0 citations
Medicine

TL;DR

A scalable machine learning framework GPCR SLM is presented, that classifies GPCRs across 86 distinct families using a lightweight transformer model optimized through knowledge distillation, and demonstrates the effectiveness of combining distilled protein language models with flexible classification frameworks for high-resolution functional annotation.

Abstract

Accurate protein family classification is essential since proteins within the same family share conserved structural domains and biochemical functions that deter mine their biological roles. G protein-coupled receptors (GPCRs) represent one of the largest and most diverse protein families in eukaryotes, serving as targets for ap proximately 35% of FDA-approved drugs. While traditional sequence alignment methods, such as BLAST, provide foundational tools for identifying homologous sequences, they exhibit limited accuracy in distinguishing closely related GPCR families with low sequence homology. Recently, deep learning approaches offer promising accuracy; however, they employ fixed-size classification architectures that force newly discovered protein families into pre-existing categories, preventing the recognition of novel families and limiting scalability as the protein universe expands. In this work, we present a scalable machine learning framework GPCR SLM, that classifies GPCRs across 86 distinct families using a lightweight transformer model optimized through knowledge distillation. Our approach achieved an overall ac curacy of 99%, significantly outperforming BLAST (86.1%) and HMMER (91%), while demonstrating substantial computational efficiency with an average speedup of 33.5× compared to large protein language models. These results demonstrate the effectiveness of combining distilled protein language models with flexible classification frameworks for high-resolution functional annotation.

View source

Similar papers

Open access Aug 2026

PLMView: collaborative protein language model representations for fast and scalable specialized protein function inference

Applications to thioredoxins, visual opsins, and Tara Oceans environmental diatom cold-shock proteins show that PLMView can move from interpretable residue-level determinants in well-studied protein families to large-scale environmental functional discovery, linking molecular specialization to ecological distribution a...

Vinh-Son Pho, Alessandro Natale Bianchi, Mattéo Scarsini et al. · 0 citations
Open access Aug 2026

Learning from human and chemical languages to predict biological function

PubCheF-1, a deep learning model that predicts literature-derived biological function directly from chemical structure, establishes that machine learning-based prediction of biological function derived from the language of scientific literature allows the identification of bioactive molecules at high hit rates, thereby...

Clayton W. Kosonocky, Nikol Kadeřábková, Kangsan Kim et al. · 0 citations
Open access Aug 2026

Data-Centric Evaluation of Protein Function Prediction Pipelines

Findings show that performance estimates in protein function prediction should be interpreted as outcomes of complete data-centric workflows rather than isolated properties of predictive models.

Nicole Soto-García, Norma Murillo-Acevedo, Julián García-Vinuesa et al. · 0 citations
Open access Sep 2026

Squidly harnesses enzyme functional hierarchy and contrastive learning to efficiently predict catalytic residues from sequence

Enzymes present a sustainable alternative to traditional chemical industries, drug synthesis, and bioremediation applications. Because catalytic residues are the key amino acids that drive enzyme function, their accurate prediction facilitates enzyme function prediction. Sequence similarity-based approaches such as BLA...

W. J. Rieger, Mikael Bodén, Frances H. Arnold et al. · 0 citations
Open access Aug 2026

DHST: A Deep Hybrid Structure–Topology Framework for Accurate Protein Function Prediction

DHST is proposed, a deep hybrid structure–topology framework that integrates sequence semantics from a pretrained protein language model with local structural information learned by a residual graph convolutional network and introduces site-specific persistent homology to encode multi-scale topological invariants and a...

Bin Lu, Fujun Xiang, Hai-Long Wang et al. · 0 citations
Open access Aug 2026

Sequence-centric deep learning druggability prediction using protein language models with multi-scale attention and feature fusion

Overall, DrugPLMFormer provides a reproducible, leakage-aware framework for retrospective sequence-based druggability screening and target prioritization, while prospective validation and experimental confirmation remain necessary before operational deployment.

Z. Kafi, Khosro Rezaee, Hossein Eslami · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.