Skip to content
Open access

TPT: a compact CNN-transformer encoder for efficient microbial small protein modeling

Aug 2026 · Frontiers in Microbiology · Vol 17 · 0 citations · 55 references
Medicine

TL;DR

TPT provides a compact and computationally efficient encoder with competitive performance for microbial peptide analysis and suggests that effective smORF representation may benefit from pretraining objectives and inductive biases tailored to short, rapidly evolving sequences, rather than from model scale alone.

Abstract

Introduction Microbial small proteins, encoded by small open reading frames (smORFs), play essential roles in antimicrobial activity, metabolic regulation, and signaling pathways. However, their short length and rapid evolutionary rate present significant challenges for computational modeling. Methods We introduce TinyProteinTransformer (TPT), a lightweight CNN-Transformer hybrid encoder pretrained on the Global Microbial smORF Catalog (GMSC; >280 million smORF families). TPT integrates multi-scale convolutional filters to capture local sequence motifs with Transformer layers for broader contextual modeling, and is trained jointly with masked language modeling and contrastive learning to capture both residue-level and sequence-level representations. We evaluated TPT on six downstream peptide/protein classification benchmarks spanning antimicrobial peptides (AMP), toxic peptides (TOX), bacteriocins (BCN), anti-CRISPR proteins (Acr), quorum-sensing peptides (QSP), and cell-penetrating peptides (CPP) using frozen-encoder linear probing. Results On the AMP and TOX tasks, TPT (103 M parameters) matched ESM2-150 M in predictive performance (AUC 0.930 vs. 0.928 for AMP; 0.930 vs. 0.925 for TOX) while achieving 4.3-fold faster inference. Across the evaluated benchmarks, TPT showed competitive performance compared with larger protein language models. Ablation experiments identified the contrastive objective as an important contributor to representation quality: its removal reduced AUC across all six downstream tasks and produced performance patterns consistent with MLM-only pretraining. Conclusion These results suggest that, within the evaluated benchmarks, effective smORF representation may benefit from pretraining objectives and inductive biases tailored to short, rapidly evolving sequences, rather than from model scale alone. TPT therefore provides a compact and computationally efficient encoder with competitive performance for microbial peptide analysis.

Read PDF

Similar papers

Open access 2026

Benchmarking CNN and LSTM Models for Genetic Mutation Classification across Diverse Sequence Encoding Techniques

This study evaluates four encoding schemes—one-hot, k-mer (substring-based encoding), embeddings, and Position-Specific Scoring Matrix (PSSM) using Convolutional Neural Networks (CNNs) and Long Short-Term Memories (LSTMs) and shows that k-mer encoding achieved the highest accuracy.

T. Kurniawan, Deshinta Arova Dewi, Randy Joy Magno Ventayen · 0 citations
Review Aug 2026

Decoding structure-function synergies in plant proteins using AlphaFold and deep learning tools.

The transition to sustainable plant-based proteins requires a molecular-level understanding of structure-function relationships that traditional analytical techniques struggle to fully characterize. This comprehensive review explores the potential of artificial intelligence in elucidating the structure-function relationship of plant proteins, particularly addressing AlphaFold and deep learning-based structural predictors to bridge this characterization gap. Through neural network architectures including graph convolutional networks, attention-based transformers, and protein language models, primary sequences are mapped to three-dimensional structures and correlated with macroscopic techno-functional attributes. Comparative profiling of soybean glycinin, pea vicilin, and faba bean legumin demonstrates how computational modeling bridges genomic sequences with food material performance. Specifically, AlphaFold structural predictions reveal that soybean glycinin G1 possesses a total monomer solvent accessible surface area (SASA) of ∼18,500 Å2, a highly hydrophobic fraction of ∼41.2% (∼7,622 Å2), and 31 β-sheet strands. This unique spatial topography, coupled with conserved disulfide bridges, specifically the mature intramolecular Cys12-Cys45 bond (precursor Cys35-Cys68) and the interchain disulfide bond involving mature acidic Cys88 (precursor Cys111) and basic Cys298 (precursor Cys321), explains its superior interfacial adsorption, film-forming capacity, and strong hydrogel network formation compared to pea vicilin, which lacks disulfide stabilization and has a lower hydrophobic fraction (∼37.5%). However, a critical limitation has been identified: prior state modeling restricts AlphaFold predictions to static native conformations under idealized conditions (pLDDT: 84.35), whereas processing-dependent technological functionalities emerge in non-equilibrium states. Thus, structural templates need to be integrated with processing parameters (pH, temperature) and molecular dynamics to simulate conformational transitions and active site exposure; this represents a paradigm shift in computer-aided precision food design.

Evrim Unal, Esra Capanoglu, Asli Can Karaca · 0 citations
Book Open access Aug 2026

MotRNA: Encoding RNA Motifs via Explicit N-gram Memory

Crucially, the analysis reveals that the backbone encoder becomes self-contained post-training: the memory module facilitates the internalization of motif semantics into the model parameters, allowing the backbone to retain performance advantages even when the memory is detached during inference.

Xiangyu Ji, Xin Wang, Yang Zhang et al. · 0 citations
Open access Jul 2026

DFN-kcr: a dual-branch deep learning model with attention-guided fusion for predicting lysine crotonylation sites in human non-histone proteins

Introduction Lysine crotonylation (Kcr) is extensively present in human non-histone proteins and plays a critical regulatory role in essential biological processes, including cell signaling and metabolic regulation. However, conventional wet-lab approaches for Kcr site identification are costly, time-consuming, and ill-suited for large-scale profiling. Although computational prediction methods have garnered increasing attention in recent years, there remains a notable lack of efficient and specialized tools tailored specifically for Kcr site prediction in human non-histone proteins. Methods To address this gap, we propose DFN-Kcr, a dual-branch deep learning model explicitly designed for human non-histone Kcr site prediction. DFN-Kcr employs a dual-input strategy combining ProteinBERT embeddings and integer-encoded amino acid sequences, leveraging residual convolutional networks to capture local sequence motifs and a Transformer architecture to model long-range contextual dependencies. A branch-level gating attention fusion mechanism is further introduced to effectively integrate complementary features from both branches. Results Extensive experiments demonstrate that DFN-Kcr significantly outperforms state-of-the-art methods, with its sensitivity (Sn), specificity (Sp), accuracy (Acc), and Matthews correlation coefficient (MCC) being 0.8244, 0.7679, 0.7961, and 0.5932, respectively. Discussion This model offers a reliable solution for high-throughput identification of Kcr sites in non-histone proteins. To facilitate community access, we have deployed a user-friendly web server at http://www.lzzzlab.top/dfnkcr/.

Xin Wei, Siqin Hu, Chen Lin · 0 citations
Open access Aug 2026

DNCLA: A Deep Learning Model for TFBS Identification Based on Structural and Conformational Properties of Nucleotides and Dinucleotides

Identifying transcription factor binding sites (TFBSs) is fundamental to understanding complex gene regulatory mechanisms and the functions of non-coding regions. Although existing methods have achieved substantial strides, capturing both local structural features and long-range spatial dependencies within DNA sequences remains a major challenge for improving prediction accuracy. In this study, we propose DNCLA, a deep learning model that synergizes multisize convolutional fusion, Bidirectional Long ShortTerm Memory (Bi-LSTM) networks, and a multi-head self-attention mechanism. At the feature extraction level, DNCLA breaks through the limitations of traditional single-sequence encoding by fusing Nucleotide Chemical Properties (NCP) with Dinucleotide Physicochemical Properties (DPCP). NCP provides a refined characterization of chemical differences between bases based on ring structures, hydrogen bond sites, and functional group properties, while DPCP introduces parameters such as local structural stability and geometric flexibility of the DNA. Subsequently, the model extracts spatial evolution from these high-dimensional features through a multi-size convolutional module; captures long-range spatial dependencies using Bi-LSTM layers; and employs a multi-head self-attention mechanism to achieve adaptive weight distribution of global features, thereby enhancing the perception of key regulatory motifs. Results from training and testing the proposed model on 165 ChIPseq datasets demonstrate that DNCLA possesses robust generalization capabilities and high predictive performance in TFBSs identification. This suggests that the incorporation of physicochemical features better elucidates the essence of interactions between transcription factors and DNA.

Jingjue Wei, Jie Feng · 0 citations
Open access Jul 2026

Transformer-based Discovery of Antimicrobial Peptides and Prediction of their Antibacterial Activity

With the worsening crisis of antimicrobial resistance, many researchers are now exploring new types of antibacterial agents, and antimicrobial peptides (AMPs) have begun to attract much attention. However, the experimental identification of AMPs in the large space of natural and synthetic sequences is still slow, expensive and labour-intensive. AMP-Transformer is a two-stage deep learning framework that combines self-supervised pre-training on large-scale unlabelled protein corpora with supervised fine-tuning on curated AMP datasets in this study. The model is a multi-layer bidirectional Transformer encoder that learns contextual residue representations via masked residue modelling, and is then fine-tuned for two coupled tasks: binary discrimination of AMPs from non-AMPs and regression of minimum inhibitory concentration (MIC) values. We used the benchmark datasets built by DBAASP v3, DRAMP 4.0 and other recently published experimental collections for evaluation. On an independent test set, AMP-Transformer had an accuracy of 95.3%, a Matthews correlation coefficient of 0.906, and an area under the receiver operating characteristic curve (AUC-ROC) of 0.986, and outperformed the support vector machine, random forest, convolutional, recurrent and hybrid baselines significantly. Ablation studies show that self-supervised pre-training and multi-head self-attention are the two main contributors, accounting for about 4% of the accuracy increase. Analysis of the learned attention maps shows that the model has independently learned the amphipathic periodicity and cationic residue enrichment characteristic of membrane-active peptides, thereby providing a degree of mechanistic interpretability that is rare among black-box predictors. A subsequent screening of metagenomic open reading frames also produced a ranked list of candidate AMPs with low sequence identity to any training examples, demonstrating the value of the framework for early-stage discovery. Based on the above results, Transformer-based protein language models are relatively stable, interpretable and scalable paradigms for AMP discovery and activity prediction, and they can help establish a practical computational pipeline that selects promising peptide candidates for experimental validation at a lower cost compared with traditional methods.

Shuwen Pan, Eason Soo, Konken Wong · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.