Jun 2026· Journal of Chemical Information and Modeling· Vol 66, pp. 6829-6836· 0 citations· 19 references
Computer ScienceMedicine
TL;DR
The Biotoxins Database (BioTD) is the largest open-source database for toxins, offering open access to 14,607 data records (8,185 activity records), covering 8,975 toxins sourced from 5,220 references and patents across over 900 species.
Abstract
Biotoxins, mainly produced by venomous animals, plants, and microorganisms, exhibit high physiological activity and unique effects such as lowering blood pressure and analgesia. A number of venom-derived drugs are already available on the market, with many more candidates currently undergoing clinical and laboratory studies. However, drug design resources related to biotoxins are insufficient, particularly because of a lack of accurate and extensive activity data. To fulfill this demand, we developed the Biotoxins Database (BioTD). BioTD is the largest open-source database for toxins, offering open access to 14,607 data records (8,185 activity records), covering 8,975 toxins sourced from 5,220 references and patents across over 900 species. The activity data in BioTD are categorized into five groups: Activity, Safety, Kinetics, Hemolysis, and other physiological indicators. Moreover, BioTD provides data on 1,532 mutants, refines the whole sequence and signal peptide sequences of toxins, and annotates disulfide-bond information. All of the data in the database can be downloaded for free. Given the importance of biotoxins and their associated data, this new database is expected to attract broad interest from diverse research fields in drug discovery. BioTD is freely accessible at http://biotoxin.net/.
Mass spectrometry (MS) is the leading technology for identifying proteins in complex biological samples. It relies on the use of tandem MS alongside a reference database of canonical protein sequences to computationally identify peptides and their parent proteins. The canonical sequences represent the most widely expressed and functionally validated forms of proteins. Consequently, disease-induced or disease-supportive variants, such as those associated with cancer, will evade detection if they are absent from the database. To address this challenge, this study introduces a revised release of the Unkown Mutation Analysis (XMAn) database by incorporating coding missense and nonsense mutations from the latest versions (v103) of the COSMIC Genome Screen Mutants (GSM) and Cancer Gene Census (CGC) datasets in two distinct FASTA-formatted peptide databases comprising 3,848,499 and 312,658 variants, respectively. The mutated peptides were matched to reviewed, non-redundant UniProt Homo sapiens protein entries (18,362 and 746), and characterized in terms of nucleotide- and amino acid mutation frequencies, peptide length distributions, and associations between specific single-nucleotide (SNV) and single amino acid (SAAVs) variants. Applied to the analysis of MDA-MB-231 breast cancer cell-membrane protein fractions, the database enabled the identification of 300+ high-quality variant peptides - several localized to functional protein-binding and catalytic domains - and 23 aberrant protein products mapped to the CGC dataset. The database is hosted and available for download on Zenodo (XMAn/gsm doi: 10.5281/zenodo.21781023; XMAn/cgc doi: 10.5281/zenodo.21781514) or can be accessed through https://sites.google.com/vt.edu/xman-db/home.
Joshua R. S. Haueis, I. Lazar· bioRxiv· 0 citations
Druggable proteins are proteins that can be specifically bound and modulated by drug molecules, with such modulation expected to produce therapeutic effects. The identification and validation of druggable proteins are central steps in modern drug discovery. With the rapid advancements in biological big data and artificial intelligence, computational methods based on bioinformatics have become essential for large-scale screening of druggable proteins. This review systematically summarizes the latest research progress in this field, providing a comprehensive analysis across three dimensions: data resources, feature engineering, and predictive models. The main contributions are 4-fold: First, we systematically describe three major types of databases: pharmacological annotation databases, quantitative affinity databases, and multi-source integrated databases, and dataset construction strategies as well as strategies for handling imbalanced samples are explored. Second, multi-level feature representations across four dimensions are categorized into sequencederived features, physicochemical properties, evolutionary information, and structural information, and the corresponding methods and tools are summarized. Third, we provide a detailed review of the evolution and application paradigms of predictive models, from classical machine learning and ensemble learning to deep learning architectures and protein large language models. Fourth, based on a systematic evaluation of representative studies, current limitations in areas such as few-shot learning and model interpretability are highlighted, and offer a forward-looking perspective on future research directions. This review aims to serve as a comprehensive reference for druggable protein prediction algorithms and to provide a solid theoretical and technical roadmap for computationally driven discovery of novel drug targets.
Hong-Qi Zhang, Hong-Ling Wang, Shang-Hua Liu et al.· Current Drug Targets· 0 citations
GPCRs make up the largest family of human membrane proteins and of drug targets. Decades of experimental and structural work have revealed how these receptors operate at the molecular level, but this work has covered only a fraction of the superfamily, leaving the majority of receptors underexplored. To address this problem, we developed the GPCR Evolution Database, gpcrevolution.org, an open-access resource that makes high-quality evolutionary analysis of GPCRs available to every laboratory. Through our database, each human receptor can be examined against its own orthologs, where per-residue conservation reveals the sites that evolution has protected throughout that receptor’s history. These orthologous lineages can then be compared with their paralogs within the same GPCR family, to distinguish residues shared across paralogous lineages, which likely support ancestral functions, from the lineage-specific residues that underlie functional differentiation between receptor subtypes. The database provides ortholog sets, multiple sequence alignments, phylogenetic trees and per-residue conservation scores for over 800 human GPCRs. The resource presents these through interactive conservation plots, sequence logos and snake plots, and provides tools for comparing conservation across two or more paralogous lineages. Overall, the GPCR Evolution Database provides the data and tools for researchers to examine GPCR function through an evolutionary lens, allowing molecular insights from well-studied receptors to be extended to their underexplored relatives.
Understanding protein mechanisms in health and disease requires characterizing the functional roles of individual amino acid residues. To explore the role of residues and their mutations, we have developed Atlantis, a database that integrates structural and functional information at the human proteome residue level. A graph database enables complex queries and the retrieval of integrated information for multiple functional analysis of protein systems. A Model Context Protocol (MCP) connector allows the interrogation of the resource through Large Language Models (LLMs) or agentic frameworks for biomedical research. Atlantis annotates over 11M residues across 20k human proteins, identifying hundreds thousands intra- and inter-protein contacts in PDB as well as AlphaFoldDB structures. We also provide the possibility to analyze and integrate predicted 3D complexes inputted by the user, and we showcased these features on hundreds of AlphaFold-multimer complexes of GPCRs and LRRK2 interaction networks. The tool is freely accessible at https://atlantis.bioinfolab.sns.it/. GRAPHICAL ABSTRACT
Natalia De Oliveira Rosa, Piergiorgio Ferronato, M. Varisco et al.· bioRxiv· 0 citations
Tuberculosis is caused by the bacterium Mycobacterium tuberculosis and is the leading cause of death from infectious diseases worldwide, being considered a granulomatous infection. The quinoline molecules were chosen because they possess antifungal and antimicrobial properties, which are normally related to their biological activities, being a privileged structure in medicinal chemistry, capable of modulating multiple targets, including kinases. The target prediction revealed a strong association with vascular endothelial growth factor receptor 2 (KDR), with 1300 and 1447 similar compounds. This article shows the structural reactivity of the four derivatives of 1,4-dihydro-4-oxo-quinoline-3-carbohydrazide, evaluated through DFT calculations in vacuum and DMSO (B3LYP/6–311 + + G(d,p)), using the ORCA 5.04 program. In addition, this study also used computational approaches of virtual screening and ADMET prediction to evaluate pharmacokinetic properties. The analyses were performed using the softwares SwissADME, ADMETlab 3.0, admetSAR 3.0, pkCSM, Pred-hERG 5.0, StopTox, and ADMET Prediction Service—LMC, and involved the evaluation of oral bioavailability (0.55 for all compounds), intestinal permeability (Caco-2: −4.708 to −4.756), toxicity (non-toxic), and pharmacokinetic profile, selecting the compounds with the best characteristics for absorption and distribution. The results showed that the QNL1 and QNL3 derivatives were the most favorable due to high intestinal absorption (> 95%) and apparent permeability (Papp > 1.0 × 10 cm/s), showing potential as a future drug. In summary, the findings show these compounds as promising candidates for the treatment of tuberculosis, E. coli bacteria, and the fungus Aspergillus fumigatus.
M. Sales, Caroline Do Nascimento Gonçalves, Abraão Lucas Silva dos Santos et al.· Discover Chemistry· 0 citations
Fermented products, such as cheeses, represent an increased interest mainly in their functional properties. The present database places particular emphasis on peptides, which are encrypted in proteins and subsequently released during proteolysis. Jihočeská Niva cheeses have been designated as Protected Geographical Indications (PGIs), recognizing their unique geographical origins and traditional craftsmanship. However, the existing scientific literature on their general characteristics remains limited. It is essential that new research expand beyond basic nutritional or technological characterization to include the evaluation of diverse functional properties in such products. This article contains data obtained from Jihočeská Niva cheeses at four ripening ages. The main data corresponds to the short-chain peptides profiles present in the non-protein nitrogen (NPN) and ethanol soluble (EtOH-SN) crude nitrogen fractions. The correlations of individual peptides with the fractions' antioxidant capacity measured by the 2,2-diphenyl-1-picrylhydrazyl (DPPH·) and 2,2′-azino-bis(3-ethylbenzothiazoline-6-sulfonic acid) (ABTS·+) assays, along with total phenolic content (TPC) were also investigated. The peptide profiles were obtained by RP-HPLC. Additionally, data on a comprehensive characterization of the cheeses is presented. This includes assessment of moisture, fat, water activity, pH, particle size, and hardness. The chromatographic analysis of cheese fractions resulted in 12 chromatograms that are placed in CSV files. These files contain information that could be used to compare new samples or to perform other statistical analyses. The data of the peptide’s integration reports are presented, as well as the individual peaks' correlation with the antioxidant capacity.
S. T. Martín-del-Campo, A. Pérez-Alva, B. R. Shah et al.· Data in Brief· 0 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduJul 22, 2026
Known for his clear and elegant writing style, Bertsekas shaped fields from control and optimization to large-scale computation and artificial intelligence.