Shifu, a framework of three coupled parts, scores a model on three axes (correctness, breadth across diverse RNAs, and whether its confidence can be trusted) rather than one number, and releases the dataset, code, and model backbones.
Abstract
Deep learning has advanced RNA secondary-structure prediction by bypassing explicit energy rules to capture long-range dependencies, yet progress is limited less by model scale than by how structures are measured: single scores hide where and why models fail, and benchmark scores can reflect memorization of one dataset rather than genuine generalization. We address this with Shifu, a framework of three coupled parts. Shifu-Corpus is a leakage-audited dataset of 254123 sequences from six databases, with family-aware splits certified free of exact and near-duplicate leaks. The Shifu Trifecta scores a model on three axes (correctness, breadth across diverse RNAs, and whether its confidence can be trusted) rather than one number. Shifu-LMR, a family of compact RNA language models, serves as controlled experiments: changing the training corpus shifts accuracy by 0.13, and a 65-million-parameter model, Shifu-LMR-Nano, leads on correctness while running on a laptop. We release the dataset, code, and model backbones.
RNA motifs and Low-Complexity Repeats (LCRs)—recurrent, conserved sequences of nucleotides—serve as the fundamental vocabulary of RNA structure and biological function. While current RNA foundation models have revolutionized sequence modeling through Transformer architectures, they predominantly prioritize modeling global dependencies within the raw sequence. To address this limitation, we propose MotRNA, a motif-aware RNA language model integrating an Explicit N-gram Memory mechanism. Unlike standard implicit neural modeling, MotRNA employs a conditional memory module to explicitly retrieve and fuse motif embeddings based on input k-mers. This mechanism provides direct access to strictly defined local patterns, complementing the global context modeled by self-attention. We validate MotRNA on large-scale RNA datasets. The model achieves a 9.6% absolute improvement in Masked Language Modeling (MLM) accuracy at a 30% masking ratio compared to state-of-the-art baselines. Notably, MotRNA maintains high robustness under extreme masking ratios, indicating effective biological signal reconstruction from sparse context. Crucially, our analysis reveals that the backbone encoder becomes self-contained post-training: the memory module facilitates the internalization of motif semantics into the model parameters, allowing the backbone to retain performance advantages even when the memory is detached during inference.
Xiang-Yu Ji, Xin Wang, Yang Zhang et al.· Proceedings of the 32nd ACM...· 0 citations
The state of the art in RNA structure prediction is reviewed, covering key training datasets, community benchmarks, and the performance of current models, as well as perspectives on integrating other data modalities, such as chemical probing signals and RNA modifications, as well as the emerging role of generative models.
Lambert Moyon, Annalisa Marsico· Current Opinion in Structura...· 0 citations
Motivation RNA language models learn representations that support structure and function prediction, but which biological concepts their hidden states encode remains unclear. Sparse autoencoders (SAEs) decompose hidden states into interpretable features, yet have not been applied to RNA language models, where byte-pair tokenization breaks the one-token-one-nucleotide correspondence that nucleotide-level attribution assumes. Results We present SPIRAL, a layer-wise SAE analysis of BiRNA-BERT. Independent SAEs at layers 0, 5, and 11 expand each 768-dimensional hidden state into 6,144 features while preserving model behaviour (explained variance above 0.99997; masked-language-model sequence recovery near 99.7%). Tokenizer-aware offset propagation aligns features to nucleotides: at layer 5, 44.3% of tested features are significantly associated with bpRNA secondary-structure classes (mean enrichment 1.61 ×), and all 1,237 eligible features with RNAcentral RNA types. Sparse profiles raise k-nearest-neighbour balanced accuracy from 0.328 to 0.359 over dense embeddings at layer 5. Availability and Implementation Source code is available at https://github.com/SadatHossain01/SPIRAL; the code, evaluation data, and trained SAE checkpoints are archived at https://doi.org/10.5281/zenodo.21891845. Contact mrahman@cse.buet.ac.bd Supplementary information Supplementary data are presented alongside the manuscript.
M. Hossain, MD. Roqunuzzaman Sojib, Md Toki Tahmid et al.· bioRxiv· 0 citations
In recent years, deep-learning has revolutionized protein structure prediction, achieving remarkable speed and accuracy. RNA structure prediction, however, has lagged behind. Although several methods have shown moderate success in predicting RNA secondary and tertiary structures, none have reached the accuracy observed with contemporary protein models. The lack of success of these RNA structure prediction models has been proposed to be due to limited high-quality structural information that can be used as training data. To probe this proposed limitation, we developed a large and diverse dataset comprising paired RNA sequences and their corresponding secondary structures. We assessed the utility of this enhanced dataset by retraining on a deep-learning model, SincFold. We find that SincFold exhibited improved performance on a set of previously unseen RNA families, enhancing its capability to predict accurate de novo RNA secondary structures. We additionally implemented Lyra-TransPred, which achieved the highest mean F1 and MCC among the evaluated models while requiring substantially less training time per epoch. The RNASSTR dataset provides a substantial advance for RNA structure modeling, laying a strong foundation for the development of future RNA secondary structure prediction algorithms.
Conner J. Langeberg, Taehan Kim, Roma Nagle et al.· RNA: A publication of the RN...· 0 citations
The results show that examining branching configurations optimal modulo the branching energy provides structural information beyond the standard MFE prediction, and the proposed energy-filtering approach yields a compact set of alternative structural hypotheses that can complement Boltzmann sampling and provide a practical source of candidate helices for downstream computational or experimental analysis.
Yuta Hozumi, Svetlana Poznanovic, Christine E. Heitsch· Biophysical Journal· 0 citations
A cheap density prior is introduced over natural protein activations and keeps only the candidates that remain typical under it, a training-free post-hoc step the authors call Mahalanobis filtering that improves both the property score and the structural plausibility of the sequences it keeps at negligible cost, without touching the generator, and transfers across different guidance methods.
Shuibai Zhang, Xin-Chi Liu, Fred Zhangzhi Peng et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.