Targeted covalent inhibition is an important strategy in modern drug discovery, with cysteine being the most common residue targeted for covalent ligands. Accurate identification of covalently ligandable cysteines is therefore essential, especially for traditionally "undruggable" targets. However, structure-based methods depend on available and reliable protein structures, while sequence-based methods remain scarce and require further improvement. Here, we present CCSite, a protein language model-based framework for discovering covalently ligandable cysteines from protein sequences. It uniquely integrates low-rank adaptation of ESM Cambrian (LoRA-ESMC) with a cysteine-centered encoder-decoder module to capture local microenvironment features and long-range contextual information. Benchmarking and independent evaluation showed that CCSite achieved competitive performance without requiring 3D structures. Moreover, a real-world application revealed that CCSite was capable of prospectively identifying experimentally validated covalent cysteines, and a large-scale screen of human pathogenic X-to-Cys mutations further identified over 2,000 neo-cysteines as covalently ligandable candidates for future covalent drug development.
Yan-Lin Ren, M. Mou, Yi-Miao Zhu et al.· Journal of Medicinal Chemist...· 0 citations
The Comprehensive VS Platform with AI Engine (CVSP-AIE) for drug discovery from compound libraries integrates three AI models: KarmaDock, a fast docking model that directly updates atomic coordinates; CarsiDock, an accurate docking model that predicts protein-ligand distances and reconstructs binding poses; and RTMScore, an accurate scoring model that learns residue-atom distance distributions for affinity prediction.