Skip to content

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access Aug 2026

A large-scale cryo-EM RNA motif dataset and benchmark for machine learning-based structure modeling.

MOTIVATION RNA functions in gene regulation, viral replication, and cellular control are tightly coupled to three-dimensional structure and local conformational features. Cryogenic electron microscopy (cryo-EM) now enables RNA structure characterization across a broad resolution range, but full maps are large, heterogeneous, and variable in local resolution. RNA secondary structural motifs, including hairpins, internal loops, and bulges, provide recurring local units for interpreting RNA density, comparing structures, and developing machine-learning models. Existing cryo-EM-based methods generally focus on complete maps, chains, residues, or atomic model construction rather than motif-level representations, partly because large-scale motif-resolved cryo-EM datasets remain limited. RESULTS We present an open-source dataset of more than 100,000 motif-resolved cryo-EM density segments paired with atomic structures, spanning 25 RNA secondary structural motif classes and resolutions from 1.5 Å to 34.0 Å. Each motif is represented as a standardized 3D voxel grid with voxel-level labels for RNA backbone, ribose sugar, and nucleobase components. Motif-level map-model agreement was evaluated using masked cross-correlation (CCmask) and atom-level Q-scores, revealing resolution-dependent trends in regional density agreement and atomic resolvability. As a baseline benchmark, a 3D convolutional neural network trained on a curated, class-balanced, primarily high-resolution subset distinguished five motif/background classes, achieving macro-averaged sensitivity of 0.836 ± 0.019, specificity of 0.958 ± 0.005, balanced accuracy of 0.897 ± 0.012, and G-mean of 0.894 ± 0.013. AVAILABILITY AND IMPLEMENTATION Source code, pipeline implementation, benchmark datasets, and an interactive web application are available at GitHub (https://github.com/DrDongSi/3DEM-RNA-Motif-Dataset), Zenodo (https://zenodo.org/communities/3dem-rna-motif-dataset), and Hugging Face Spaces (https://huggingface.co/spaces/houlab/arsma-cryoem).

Chandramathi Murugadass, Hajira Rana, Brent M. Znosko et al. · 0 citations
Open access Jul 2026

Enhancing feature selection for ordinal outcomes using resampling-based sparse linear discriminant analysis

Abstract Motivation High-dimensional biomedical datasets with ordinal outcomes—such as cancer stages or treatment responses—pose significant challenges for feature selection due to strong predictor correlations and limited sample sizes. Sparse Linear Discriminant Analysis (sLDA) is widely used for simultaneous classification and feature selection. However, concerns about model stability and reproducible feature selection persist, particularly in the presence of pronounced collinearity inherent in biomedical data. Consequently, direct application of sLDA often fails to capture a reproducible set of biologically coordinated markers, resulting in signatures that lack robustness and interpretability. Results We propose a resampling-based ensemble sLDA framework that integrates bootstrapping and subsampling to improve the stability of feature selection. By aggregating results across multiple resampled datasets, the method identifies features based on Variable Inclusion Probability (VIP) rather than relying on coefficients from standard sLDA. Compared with standard sLDA, this ensemble strategy reduces sensitivity to data perturbation and improves the stability and reproducibility of selected feature sets. Simulation studies demonstrate that the proposed ensemble framework achieves more accurate and consistent recovery of ground-truth predictors compared with the standard (non-resampled) sLDA. Applications to kidney renal papillary cell carcinoma staging and glioma grading datasets further suggest that this framework can improve predictive performance and identify biologically interpretable and reproducible feature sets, highlighting its potential utility for reliable biomarker discovery in precision medicine. Availability and implementation The source code used in the study is available via GitHub at https://github.com/ryan-wng/RE-sLDA

Yin Liu, Ryan Wang, Dong Si · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.