Skip to content

CABA-Bind: Confounder-aligned backdoor adjustment for debiased RNA-ligand binding prediction.

Aug 2026 · Computational biology and chemistry · Vol 125, pp. 109293 · 0 citations · 48 references
Medicine

TL;DR

CABA-Bind improves the reliability and ligand-specific interpretability of RNA-ligand molecular recognition modeling, and provided computational evidence suggesting that model-highlighted RNA regions around G17 and A53/A54 may contribute to ligand-associated recognition.

Abstract

Context

RNA-ligand molecular recognition is important for RNA-targeted drug discovery and candidate small-molecule prioritization. However, RNA-ligand binding datasets often contain biased associations between RNA sequences and binding labels, which may cause models to rely on sequence-driven cues rather than ligand-dependent binding features. Such sequence-bias-driven reliance can reduce model reliability in challenging prediction scenarios, especially when evaluating unseen RNA targets, structurally dissimilar ligands, or hard decoy molecules.

Methods

We propose CABA-Bind, a causal debiasing framework for RNA-ligand binding prediction. CABA-Bind encodes RNA sequences and ligand SMILES using RNA-FM and ChemBERTa, constructs confounder centers by K-means clustering to represent recurrent RNA prior patterns, and combines RNA-Confounder Alignment with backdoor-adjusted prediction to reduce the influence of RNA sequence-driven bias. The model was evaluated on the Robin and Biosensor datasets using four data-splitting strategies, prior-dependency metrics, ablation studies, hard decoy ranking, and structure-guided interpretation. CABA-Bind reduced the Score-prior |ρ| by 40.8% compared with the baseline model, achieved an MRR of 0.65 in hard decoy evaluation, and provided computational evidence suggesting that model-highlighted RNA regions around G17 and A53/A54 may contribute to ligand-associated recognition. These results suggest that CABA-Bind improves the reliability and ligand-specific interpretability of RNA-ligand molecular recognition modeling.

View source

Similar papers

Aug 2026

CoCoBind: Consistency-Contrastive Multitask Learning for RNA–Ligand Interaction and Binding Site Prediction

CoBind is presented, a multitask deep learning framework that jointly predicts RNA–compound interactions and nucleotide-level binding-site probabilities within a unified architecture and provides complementary nucleotide-level binding-site localization, supporting a site-aware view of RNA–ligand recognition under distribution shift.

Shihang Wang, Lin Wang, Wei Zhao et al. · 0 citations
Open access Aug 2026

M2-PRNet: multi-scale and multi-modal learning for protein–RNA binding affinity prediction

The results demonstrate the effectiveness of integrating multi-scale and multi-modal representations with cross-scale alignment for protein–RNA affinity prediction, and suggest that M2-PRNet can highlight relevant RNA-binding regions and support preliminary discrimination between strong and weak binders when plausible complex structures are available.

Junkai Wang, G. Luo, Yun-Song Yang et al. · 0 citations
Open access Jul 2026

Adaptive 2.5D base-pairing subgraph search detects RNA small-molecule binding sites

Ribo-LENS turns coarse base-pairing structure into a practical entry point for screening the vast, largely unexplored RNA target space, and depends far less on sequence homology than competing predictors.

David Nitchi, J. Waldispühl, C. Oliver · 0 citations
Open access Aug 2026

Boltz-Perturb: Improving Diversity and Accuracy in Protein-Ligand Co-Folding through Training-Free Conditioning Perturbation

Boltz-Perturb is presented, a framework for addressing small molecule binding poses through perturbing model conditioning signals during model inference, and it is demonstrated that inference-time perturbations can unlock latent structural diversity in generative co-folding models and improve protein-ligand predictions without costly retraining.

Hyeyun Jung, BoRam Lee, Alan C. Cheng · 0 citations
Open access Jul 2026

Analysing open-source protein folding models for nanobody binding prediction

These findings provide practical guidance for integrating open-source protein structure prediction models into AI-driven nanobody discovery pipelines while highlighting the need for improved generalization across antigens.

Yannick Vogt, Rebekka Roßberg, Jan Habermann et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.