Aug 2026· Interdisciplinary Sciences Computational Life Sciences· 0 citations· 51 references
Medicine
TL;DR
MBPBERT provides a scalable and efficient in silico solution for high-throughput discovery of novel MBPs and screening of peptides with metal-specific binding preferences, potentially reducing the reliance on resource-intensive experimental validation.
BiteNetI is a structure-based deep learning model that uses 3D convolutional neural networks to simultaneously localize ion-binding centers and predict binding residues for 14 biologically relevant ions, supporting comprehensive and large-scale annotation of protein-ion interactions.
Igor Kozlovskii, Petr Popov· Communications Biology· 0 citations
A Grouped Multi-Task Learning (GMTL) strategy is implemented, allowing the model to capture shared binding patterns among ligands with similar biological significance, allowing the model to capture shared binding patterns among ligands with similar biological significance.
Abstract Motivation Accurate identification of DNA-binding proteins (DBPs) and RNA-binding proteins (RBPs) is critical for elucidating transcriptional and post-transcriptional regulatory mechanisms. However, existing computational approaches often rely on inferred labels or domain-specific annotations, which limit the subsequent generalizability. Results This study aimed to introduce transformer-based classifiers for human DBPs and RBPs that rely solely on protein sequence information without engineered features or domain constraints. The models were implemented using ESM-2 with low-rank adaptation (LoRA) fine-tuning and trained on experimentally validated datasets, including chromatin immunoprecipitation sequencing (ChIP-seq) annotations for DBPs and eCLIP annotations for RBPs. Next, to evaluate biological relevance, we computed value-aware attention (VAT) scores aggregated across transformer layers to interpret model focus. In 20-fold cross-validation, the DBP model achieved an area under the receiver operating characteristic curve (AUROC) of 0.84 with a Matthews correlation coefficient (MCC) of 0.40, while the RBP model achieved an AUROC of 0.92 with an MCC of 0.46. Proteins predicted as nucleic acid-binding were enriched for known binding domains, and inspection of attention distributions revealed preferential focus on annotated functional regions rather than non-binding segments. These results demonstrate that attention-based protein language models can accurately identify nucleic acid-binding proteins directly from sequence data. Moreover, these models reveal biologically meaningful sequence determinants of binding, establishing an interpretable and scalable framework for proteome-wide characterization of protein–nucleic acid interactions. Availability and implementation Code is available on GitHub (https://github.com/CSB-hub/DRBP).
Hanjin Kim, Sung-Gwon Lee, Joo-Seong Oh et al.· Bioinformatics Advances· 0 citations
Testing the ability of common large language models to consider design principles to generate de novo proteins that bind metals and lipophilic small molecules without copying existing sequences highlights the utility of LLMs in making protein design more comprehensible and accessible to users without sophisticated design expertise.
Machine learning methods for predicting the electron ionization mass spectra from molecular structures have shown promise for environmental chemical identification, but their performance under domain-specific data scarcity remains poorly understood. We systematically compare a conventional multilayer perceptron model (NEIMS) with a Transformer-based chemical foundation model (MolFormer-XL) for the electron ionization mass spectrometry spectrum prediction under controlled few-shot conditions. Using fluorine-containing molecules as a broader proxy domain, including a PFAS-like subset, motivated by the practical challenge of detecting novel fluorinated contaminants with limited reference data, we vary the number of domain-specific training examples from 5 to 175 while maintaining fixed validation and test sets. Across all few-shot conditions and three of four evaluation metrics (weighted cosine similarity, intensity-weighted precision, and top-10 precision), MolFormer-XL consistently outperforms NEIMS, while intensity-weighted recall remains comparable between the two models. The largest performance gaps are observed in extreme data-scarcity regimes. These results demonstrate that MolFormer-XL, which combines pre-trained molecular representations with a Transformer-based architecture and learned SMILES embeddings, provides a promising approach for transfer under severe domain-specific data scarcity in environmental mass spectrometry.
Peptide-protein affinity models are often evaluated with a single data split, obscuring whether they interpolate among measurements for observed targets or generalize across peptide or target shifts. We integrated three sources of quantitative peptide-protein binding data to obtain 11,349 deduplicated pairs and benchmarked ten peptide representations, ESM-2 protein embeddings, and six regressors under peptide-similarity, within-target, and leave-target-out partitions. Across 60 matched representation-regressor configurations, mean test Spearman correlations were 0.462, 0.669, and 0.530, respectively. The top configuration shifted from ECFP-16 count fingerprints with random forest in the first two settings to HELM-BERT with Extra Trees when exact target sequences were excluded. Representation-rank correlations ranged from -0.042 to 0.624 across partitions, whereas regressor-rank correlations ranged from 0.771 to 0.943. Learning curves showed that representation differences were largest with limited supervision and narrowed as training data increased. PeptideCLM-2 adaptation and simple element-wise interaction features provided no consistent gain over a frozen encoder and direct concatenation under the tested protocols. These conclusions are specific to a dataset that pools transformed Kd, Ki, and IC50 measurements and to target exclusion at the exact-sequence level. Peptide-protein affinity benchmarks should therefore align data partitions with the intended use and jointly assess the effects of data scale, molecular representation, and downstream learner.
Jialin Tian, Darren An, Jun Li· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.