Jul 2026· Current Opinion in Structural Biology· Vol 100, pp.
103339
· 0 citations· 63 references
Medicine
TL;DR
The state of the art in RNA structure prediction is reviewed, covering key training datasets, community benchmarks, and the performance of current models, as well as perspectives on integrating other data modalities, such as chemical probing signals and RNA modifications, as well as the emerging role of generative models.
Abstract
RNA structure is central to the function of every RNA class yet the gap between annotated sequences and experimentally determined structures remains large. Computational methods to fill this gap have evolved from thermodynamic free energy minimization through supervised deep learning to self-supervised RNA language models trained on millions of sequences, progressively improving structure prediction. Here we review the state of the art in RNA structure prediction, covering key training datasets, community benchmarks, and the performance of current models. We further discuss perspectives on integrating other data modalities, such as chemical probing signals and RNA modifications, as well as the emerging role of generative models. Challenges in generalization, handling of noncanonical interactions, and contextual structure prediction remain open frontiers for the field.
These approaches improve generalisability, reduce reliance on deep evolutionary information, and enable proteome-scale prediction of RNA-binding residues, providing a route to map and interpret the molecular logic of protein-RNA interactions.
Rozeena Arif, Alfredo Castello· Current Opinion in Structura...· 0 citations
In recent years, deep-learning has revolutionized protein structure prediction, achieving remarkable speed and accuracy. RNA structure prediction, however, has lagged behind. Although several methods have shown moderate success in predicting RNA secondary and tertiary structures, none have reached the accuracy observed with contemporary protein models. The lack of success of these RNA structure prediction models has been proposed to be due to limited high-quality structural information that can be used as training data. To probe this proposed limitation, we developed a large and diverse dataset comprising paired RNA sequences and their corresponding secondary structures. We assessed the utility of this enhanced dataset by retraining on a deep-learning model, SincFold. We find that SincFold exhibited improved performance on a set of previously unseen RNA families, enhancing its capability to predict accurate de novo RNA secondary structures. We additionally implemented Lyra-TransPred, which achieved the highest mean F1 and MCC among the evaluated models while requiring substantially less training time per epoch. The RNASSTR dataset provides a substantial advance for RNA structure modeling, laying a strong foundation for the development of future RNA secondary structure prediction algorithms.
Conner J. Langeberg, Taehan Kim, Roma Nagle et al.· RNA: A publication of the RN...· 0 citations
Much of the human genome’s non-protein-coding fraction acts directly through RNA, yet the structural and functional roles encoded in these sequences remain poorly understood. Applying deep learning is hindered by scarce RNA structural data and it remains unclear what biological constraints such models can recover directly from the abundant RNA sequences alone. Here, to address these challenges, we developed NucleicBERT, a self-supervised masked-language model that learns contextual representations from single sequences without evolutionary information. Explainable artificial intelligence analyses show that the model organizes RNA sequences in latent space and encodes structural properties indicating that biologically meaningful constraints are learned from sequence correlations alone. When fine-tuned for downstream structural and functional tasks, NucleicBERT requires only single sequences while matching or exceeding current RNA prediction models. This alignment-free framework addresses the scarcity of annotated 3D RNA data while providing a rapid, computational complement to experimental techniques. By bridging abundant unlabelled sequence data with scarce structural annotations, NucleicBERT advances RNA structure prediction and informs how large language models encode biological information.
Utkarsh Upadhyay, Julian Herold, Markus Götz et al.· Nature Machine Intelligence· 0 citations
Shifu, a framework of three coupled parts, scores a model on three axes (correctness, breadth across diverse RNAs, and whether its confidence can be trusted) rather than one number, and releases the dataset, code, and model backbones.
Gabriel Galvez, Quentin Vicens· bioRxiv· 0 citations
Parnet is a multi-task foundation model trained end-to-end on 223 eCLIP-seq experiments spanning 150 RBPs to predict base-resolution RBP binding profiles directly from RNA sequence, and establishes the RBP interactome as a compact, functionally sufficient, and interpretable basis for foundation model pretraining in RNA biology.
RNA motifs and Low-Complexity Repeats (LCRs)—recurrent, conserved sequences of nucleotides—serve as the fundamental vocabulary of RNA structure and biological function. While current RNA foundation models have revolutionized sequence modeling through Transformer architectures, they predominantly prioritize modeling global dependencies within the raw sequence. To address this limitation, we propose MotRNA, a motif-aware RNA language model integrating an Explicit N-gram Memory mechanism. Unlike standard implicit neural modeling, MotRNA employs a conditional memory module to explicitly retrieve and fuse motif embeddings based on input k-mers. This mechanism provides direct access to strictly defined local patterns, complementing the global context modeled by self-attention. We validate MotRNA on large-scale RNA datasets. The model achieves a 9.6% absolute improvement in Masked Language Modeling (MLM) accuracy at a 30% masking ratio compared to state-of-the-art baselines. Notably, MotRNA maintains high robustness under extreme masking ratios, indicating effective biological signal reconstruction from sparse context. Crucially, our analysis reveals that the backbone encoder becomes self-contained post-training: the memory module facilitates the internalization of motif semantics into the model parameters, allowing the backbone to retain performance advantages even when the memory is detached during inference.
Xiang-Yu Ji, Xin Wang, Yang Zhang et al.· Proceedings of the 32nd ACM...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.