Jul 2026· International Journal of Biological Macromolecules· Vol 379, pp.
153833
· 0 citations· 31 references
Medicine
TL;DR
VAERNAGen, a novel variational autoencoder-based framework for de novo generation of RNA family sequences, outperformed the current state-of-the-art grammar-based method and is established as a state-of-the-art method for de novo generation of RNA family sequences.
Abstract
RNA plays a pivotal role in diverse cellular processes, and the rational design of functional RNA sequences is central to advancing RNA engineering. While deep generative models have shown significant promise, their generated sequences frequently lack the structural accuracy and evolutionary fidelity required for biological functionality. To address this challenge, we introduced VAERNAGen, a novel variational autoencoder-based framework for de novo generation of RNA family sequences. VAERNAGen's core innovation is a joint representation that encode aligned nucleotide sequences and their secondary structures into an 11-channel L × L two-dimensional matrix. This image-like format enables 2D convolutional neural networks to effectively learn the spatial interplay between sequence conservation and structural variation. Evaluated on two canonical Rfam families-RF00001 (5S ribosomal RNA) and RF00005 (tRNA), VAERNAGen outperformed the current state-of-the-art grammar-based method, achieving significantly higher median bit scores (105.25 vs. 89.34 for RF00001; 58.24 vs. 50.02 for RF00005). Generated sequences also exhibited greater nucleotide-level similarity to natural seed alignments and lower variance, reflecting enhanced biological plausibility and reproducibility. Together, these results establish VAERNAGen as a state-of-the-art method for de novo generation of RNA family sequences.
Background RNA inverse folding designs nucleotide sequences expected to adopt a prescribed secondary structure. Search-based solvers can optimize folding-model objectives effectively, but difficult targets can require extensive sampling, and structural optimization alone does not explicitly preserve the sequence distributions or conserved motifs of natural RNA families. Methods We developed RIFT-VAE, a Transformer-based conditional variational autoencoder that receives a context-free grammar parse-tree representation of a target secondary structure and generates nucleotide labels on the corresponding tree. The framework combines progressively richer grammar rules, self-refinement learning from generated structure-sequence pairs, and cross-entropy-method optimization in the learned latent space. We evaluated RNAfold minimum-free-energy agreement on an RNAcentral-derived test set and the EteRNA100 benchmark, compared the method with four search-based solvers under matched total time budgets, and examined GC-content control and covariance-model family annotation. Results The complete pipeline achieved RNAfold-Correct/RNAfold-MCC values of 0.833/0.994 on the RNAcentral-derived test set and 0.760/0.977 on EteRNA100. Latent-space optimization accounted for the largest increase in exact structural recovery. Under a 3,600-s total budget on EteRNA100, sequences generated by RIFT-VAE improved the exact-match rate of every tested downstream search method when used as warm starts; the largest change was observed for RNAInverse (Correct, 0.297 to 0.803; MCC, 0.505 to 0.985). The pretrained model also produced sequences with measurable correct-family covariance-model hits and supported explicit GC-content conditioning. Conclusions RIFT-VAE is best interpreted as a hybrid generative-search framework: pretraining supplies a structure- and family-informed proposal distribution, whereas latent optimization concentrates evaluations in high-scoring regions. The reported structural scores are specific to RNAfold minimum-free-energy validation and do not establish biochemical function. Orthogonal folding predictors, stricter homology-controlled splits, diversity-aware evaluation, architecture-matched dot-bracket ablations, and experimental assays remain priorities for validation.
Abstract RNA secondary structure is essential for understanding the functions of non-coding RNAs, ribosomal RNAs, and viral genomes. However, accurate prediction of long RNA structures remains challenging due to complex long-range interactions and the limited availability of long-RNA training data. We present UFold-X, a dual-branch deep learning framework that combines a convolutional encoder for local structure modeling with a Mamba-based Visual State Space Module for capturing long-range dependencies. A dynamic gating mechanism adaptively integrates the two branches according to sequence length. UFold-X was evaluated on multiple benchmark datasets containing RNAs up to 5000 nucleotides. To rigorously assess generalization, we introduced a cross-clan benchmark for long RNAs. Under this stringent setting, UFold-X achieved performance comparable to state-of-the-art classical approaches while achieving the best performance among deep learning-based methods. Additional cross-family and within-family evaluations further demonstrated robust transferability and competitive predictive performance. UFold-X also maintained excellent computational efficiency, requiring only 0.08 s per sequence on average. To assess biological consistency, we developed a SHAPE-based reactivity prediction variant (UFold-X-R) and an integrated metric, the Hybrid Reactivity-Pairing Score (HRPS). UFold-X-R showed strong agreement with experimental icSHAPE data and achieved the highest HRPS among all evaluated methods. A user-friendly web server is available at https://ufold-x.ai4bread.com.
Lai-Yi Fu, Jia-Chun Li, Rui-Qi Wang et al.· Nucleic Acids Research· 0 citations
Variational autoencoders (VAEs) trained on multiple sequence alignments (MSAs) have emerged as powerful generative models for biological sequences, with applications ranging from disease variant prediction to functional RNA design. However, standard biological VAE training treats all sequences as exchangeable, ignoring the rich evolutionary structure that organizes homologous sequences from evolutionarily close to highly divergent. We propose Evolutionary Curriculum Learning (ECL), a training strategy that exploits this structure by progressively exposing the model to sequences of increasing evolutionary distance from sampled anchors, following a power-law expansion schedule. Applied to two architecturally distinct VAE models and two biological domains--protein variant effect prediction with EVE and RNA family sequence generation with RfamGen--ECL improves downstream task performance across five random seeds per configuration. Mean ClinVar classification AUROC rises from 0.981 to 0.989 for p53; for PTEN, ECL attains 1.000 in every seed whereas the baseline is unstable (mean 0.905, falling as low as 0.54). For RNA, ECL raises mean covariance-model bit scores on all three families tested and exceeds its seed-matched baseline in 12 of 15 training runs, though with only three families the effect cannot be established as significant at the family level. Ablation experiments show that progressively expanding the sampled sequences by evolutionary distance outperforms fixed-size neighborhood sampling in addition to uniform random sampling. Evolutionary distance is therefore a useful inductive bias for ordering the training curriculum in biological sequence modeling.
Much of the human genome’s non-protein-coding fraction acts directly through RNA, yet the structural and functional roles encoded in these sequences remain poorly understood. Applying deep learning is hindered by scarce RNA structural data and it remains unclear what biological constraints such models can recover directly from the abundant RNA sequences alone. Here, to address these challenges, we developed NucleicBERT, a self-supervised masked-language model that learns contextual representations from single sequences without evolutionary information. Explainable artificial intelligence analyses show that the model organizes RNA sequences in latent space and encodes structural properties indicating that biologically meaningful constraints are learned from sequence correlations alone. When fine-tuned for downstream structural and functional tasks, NucleicBERT requires only single sequences while matching or exceeding current RNA prediction models. This alignment-free framework addresses the scarcity of annotated 3D RNA data while providing a rapid, computational complement to experimental techniques. By bridging abundant unlabelled sequence data with scarce structural annotations, NucleicBERT advances RNA structure prediction and informs how large language models encode biological information.
Utkarsh Upadhyay, Julian Herold, Markus Götz et al.· Nature Machine Intelligence· 0 citations
Protein-RNA interactions regulate diverse biological processes and are increasingly exploited in therapeutic RNA discovery, but accurate inferences of nucleotide preferences and reliable structure prediction remain challenging. Here, we present PRIS, a unified structure-based deep-learning framework that combines two complementary components: PRISeq for nucleotide probability estimation at each RNA position and PRIScore for residue-nucleotide distance prediction to discriminate native-like from incorrect poses. Both share a feature extractor that integrates an Anti-Symmetric Graph Attention Network (A-GAT) with sparse k-Maximum Inner Product (k-MPI) attention to capture long-range interactions across large graphs. PRIScore improves the selection of native-like protein-RNA predictions generated by AlphaFold3, achieving a top-1 success rate of 81.91% on a docking benchmark, compared to 79.26% for AlphaFold3. The selected structures are then fed into PRISeq, which infers position-specific binding preferences and screens RNA libraries. On a PWM benchmark, PRISeq achieved a mean absolute error (MAE) of 0.75, outperforming FoldX, Rosetta-based scoring functions, and NA-MPNN. In virtual screening against MS2 protein, PRISeq screens 129,248 RNA hairpins within 11.95 seconds, achieving the highest EF0.5% of 14.40, approximately double the best baseline. PRIS also effectively enriches active aptamers against NELF-E and GFP while preserving sequence diversity. By integrating structure selection with binding-preference inference, PRIS provides an efficient framework for large-scale RNA library screening and aptamer design.
Yi-Hao Zhao, Jing Han, Ji-Ke Wang et al.· bioRxiv· 0 citations
RNA motifs and Low-Complexity Repeats (LCRs)—recurrent, conserved sequences of nucleotides—serve as the fundamental vocabulary of RNA structure and biological function. While current RNA foundation models have revolutionized sequence modeling through Transformer architectures, they predominantly prioritize modeling global dependencies within the raw sequence. To address this limitation, we propose MotRNA, a motif-aware RNA language model integrating an Explicit N-gram Memory mechanism. Unlike standard implicit neural modeling, MotRNA employs a conditional memory module to explicitly retrieve and fuse motif embeddings based on input k-mers. This mechanism provides direct access to strictly defined local patterns, complementing the global context modeled by self-attention. We validate MotRNA on large-scale RNA datasets. The model achieves a 9.6% absolute improvement in Masked Language Modeling (MLM) accuracy at a 30% masking ratio compared to state-of-the-art baselines. Notably, MotRNA maintains high robustness under extreme masking ratios, indicating effective biological signal reconstruction from sparse context. Crucially, our analysis reveals that the backbone encoder becomes self-contained post-training: the memory module facilitates the internalization of motif semantics into the model parameters, allowing the backbone to retain performance advantages even when the memory is detached during inference.
Xiang-Yu Ji, Xin Wang, Yang Zhang et al.· Proceedings of the 32nd ACM...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.