Skip to content
Review Open access

Machine learning and language models for RNA structure prediction: Progress and perspectives.

Jul 2026 · Current Opinion in Structural Biology · Vol 100, pp. 103339 · 0 citations · 63 references
Medicine

TL;DR

The state of the art in RNA structure prediction is reviewed, covering key training datasets, community benchmarks, and the performance of current models, as well as perspectives on integrating other data modalities, such as chemical probing signals and RNA modifications, as well as the emerging role of generative models.

Abstract

RNA structure is central to the function of every RNA class yet the gap between annotated sequences and experimentally determined structures remains large. Computational methods to fill this gap have evolved from thermodynamic free energy minimization through supervised deep learning to self-supervised RNA language models trained on millions of sequences, progressively improving structure prediction. Here we review the state of the art in RNA structure prediction, covering key training datasets, community benchmarks, and the performance of current models. We further discuss perspectives on integrating other data modalities, such as chemical probing signals and RNA modifications, as well as the emerging role of generative models. Challenges in generalization, handling of noncanonical interactions, and contextual structure prediction remain open frontiers for the field.

Read PDF

Similar papers

Review Open access Aug 2026

A new dimension in protein-RNA interface prediction: Integrating protein language models and geometric deep learning.

These approaches improve generalisability, reduce reliance on deep evolutionary information, and enable proteome-scale prediction of RNA-binding residues, providing a route to map and interpret the molecular logic of protein-RNA interactions.

Rozeena Arif, Alfredo Castello · 0 citations
Open access Sep 2026

Improving RNA Secondary Structure Prediction Through Expanded Training Data.

In recent years, deep-learning has revolutionized protein structure prediction, achieving remarkable speed and accuracy. RNA structure prediction, however, has lagged behind. Although several methods have shown moderate success in predicting RNA secondary and tertiary structures, none have reached the accuracy observed with contemporary protein models. The lack of success of these RNA structure prediction models has been proposed to be due to limited high-quality structural information that can be used as training data. To probe this proposed limitation, we developed a large and diverse dataset comprising paired RNA sequences and their corresponding secondary structures. We assessed the utility of this enhanced dataset by retraining on a deep-learning model, SincFold. We find that SincFold exhibited improved performance on a set of previously unseen RNA families, enhancing its capability to predict accurate de novo RNA secondary structures. We additionally implemented Lyra-TransPred, which achieved the highest mean F1 and MCC among the evaluated models while requiring substantially less training time per epoch. The RNASSTR dataset provides a substantial advance for RNA structure modeling, laying a strong foundation for the development of future RNA secondary structure prediction algorithms.

Conner J. Langeberg, Taehan Kim, Roma Nagle et al. · 0 citations
Open access Sep 2026

NucleicBERT interprets RNA sequence space through self-supervised language modelling

Much of the human genome’s non-protein-coding fraction acts directly through RNA, yet the structural and functional roles encoded in these sequences remain poorly understood. Applying deep learning is hindered by scarce RNA structural data and it remains unclear what biological constraints such models can recover directly from the abundant RNA sequences alone. Here, to address these challenges, we developed NucleicBERT, a self-supervised masked-language model that learns contextual representations from single sequences without evolutionary information. Explainable artificial intelligence analyses show that the model organizes RNA sequences in latent space and encodes structural properties indicating that biologically meaningful constraints are learned from sequence correlations alone. When fine-tuned for downstream structural and functional tasks, NucleicBERT requires only single sequences while matching or exceeding current RNA prediction models. This alignment-free framework addresses the scarcity of annotated 3D RNA data while providing a rapid, computational complement to experimental techniques. By bridging abundant unlabelled sequence data with scarce structural annotations, NucleicBERT advances RNA structure prediction and informs how large language models encode biological information.

Utkarsh Upadhyay, Julian Herold, Markus Götz et al. · 0 citations
Open access Jul 2026

Shifu: an integrated framework for deep learning of RNA secondary structure

Shifu, a framework of three coupled parts, scores a model on three axes (correctness, breadth across diverse RNAs, and whether its confidence can be trusted) rather than one number, and releases the dataset, code, and model backbones.

Gabriel Galvez, Quentin Vicens · 0 citations
Open access Aug 2026

PARNET: A CLIP-SEQ-BASED FOUNDATION MODEL FOR RNA SEQUENCE REPRESENTATION LEARNING

Parnet is a multi-task foundation model trained end-to-end on 223 eCLIP-seq experiments spanning 150 RBPs to predict base-resolution RBP binding profiles directly from RNA sequence, and establishes the RBP interactome as a compact, functionally sufficient, and interpretable basis for foundation model pretraining in RNA biology.

Lambert Moyon, Andreina Tirabassi, Artem Baranowskii et al. · 0 citations
Book Open access Aug 2026

MotRNA: Encoding RNA Motifs via Explicit N-gram Memory

RNA motifs and Low-Complexity Repeats (LCRs)—recurrent, conserved sequences of nucleotides—serve as the fundamental vocabulary of RNA structure and biological function. While current RNA foundation models have revolutionized sequence modeling through Transformer architectures, they predominantly prioritize modeling global dependencies within the raw sequence. To address this limitation, we propose MotRNA, a motif-aware RNA language model integrating an Explicit N-gram Memory mechanism. Unlike standard implicit neural modeling, MotRNA employs a conditional memory module to explicitly retrieve and fuse motif embeddings based on input k-mers. This mechanism provides direct access to strictly defined local patterns, complementing the global context modeled by self-attention. We validate MotRNA on large-scale RNA datasets. The model achieves a 9.6% absolute improvement in Masked Language Modeling (MLM) accuracy at a 30% masking ratio compared to state-of-the-art baselines. Notably, MotRNA maintains high robustness under extreme masking ratios, indicating effective biological signal reconstruction from sparse context. Crucially, our analysis reveals that the backbone encoder becomes self-contained post-training: the memory module facilitates the internalization of motif semantics into the model parameters, allowing the backbone to retain performance advantages even when the memory is detached during inference.

Xiang-Yu Ji, Xin Wang, Yang Zhang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.