Aug 2026· Computational biology and chemistry· Vol 125, pp.
109293
· 0 citations· 48 references
Medicine
TL;DR
CABA-Bind improves the reliability and ligand-specific interpretability of RNA-ligand molecular recognition modeling, and provided computational evidence suggesting that model-highlighted RNA regions around G17 and A53/A54 may contribute to ligand-associated recognition.
Abstract
Context
RNA-ligand molecular recognition is important for RNA-targeted drug discovery and candidate small-molecule prioritization. However, RNA-ligand binding datasets often contain biased associations between RNA sequences and binding labels, which may cause models to rely on sequence-driven cues rather than ligand-dependent binding features. Such sequence-bias-driven reliance can reduce model reliability in challenging prediction scenarios, especially when evaluating unseen RNA targets, structurally dissimilar ligands, or hard decoy molecules.
Methods
We propose CABA-Bind, a causal debiasing framework for RNA-ligand binding prediction. CABA-Bind encodes RNA sequences and ligand SMILES using RNA-FM and ChemBERTa, constructs confounder centers by K-means clustering to represent recurrent RNA prior patterns, and combines RNA-Confounder Alignment with backdoor-adjusted prediction to reduce the influence of RNA sequence-driven bias. The model was evaluated on the Robin and Biosensor datasets using four data-splitting strategies, prior-dependency metrics, ablation studies, hard decoy ranking, and structure-guided interpretation. CABA-Bind reduced the Score-prior |ρ| by 40.8% compared with the baseline model, achieved an MRR of 0.65 in hard decoy evaluation, and provided computational evidence suggesting that model-highlighted RNA regions around G17 and A53/A54 may contribute to ligand-associated recognition. These results suggest that CABA-Bind improves the reliability and ligand-specific interpretability of RNA-ligand molecular recognition modeling.
CoBind is presented, a multitask deep learning framework that jointly predicts RNA–compound interactions and nucleotide-level binding-site probabilities within a unified architecture and provides complementary nucleotide-level binding-site localization, supporting a site-aware view of RNA–ligand recognition under distribution shift.
Shihang Wang, Lin Wang, Wei Zhao et al.· Journal of Medicinal Chemist...· 0 citations
The results demonstrate the effectiveness of integrating multi-scale and multi-modal representations with cross-scale alignment for protein–RNA affinity prediction, and suggest that M2-PRNet can highlight relevant RNA-binding regions and support preliminary discrimination between strong and weak binders when plausible complex structures are available.
Junkai Wang, G. Luo, Yun-Song Yang et al.· Bioinformatics· 0 citations
Ribo-LENS turns coarse base-pairing structure into a practical entry point for screening the vast, largely unexplored RNA target space, and depends far less on sequence homology than competing predictors.
David Nitchi, J. Waldispühl, C. Oliver· bioRxiv· 0 citations
Boltz-Perturb is presented, a framework for addressing small molecule binding poses through perturbing model conditioning signals during model inference, and it is demonstrated that inference-time perturbations can unlock latent structural diversity in generative co-folding models and improve protein-ligand predictions without costly retraining.
Hyeyun Jung, BoRam Lee, Alan C. Cheng· bioRxiv· 0 citations
These findings provide practical guidance for integrating open-source protein structure prediction models into AI-driven nanobody discovery pipelines while highlighting the need for improved generalization across antigens.
Yannick Vogt, Rebekka Roßberg, Jan Habermann et al.· Frontiers in Bioinformatics· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.