Skip to content
Open access

Multimodal contrastive learning for integrating molecular representations and cellular phenotypes in drug-target interaction prediction

Aug 2026 · Bioinformatics · Vol 42 · 0 citations · 32 references
Medicine

TL;DR

A two-stage contrastive learning framework integrating drug structures, protein sequences, and Cell Painting morphological profiles into a unified embedding space, which reveals pathway-specific morphological signatures associated with drug targets, providing biologically interpretable insights into drug mechanisms.

Abstract

Abstract Motivation Accurate prediction of drug-target interactions (DTIs) is fundamental to drug discovery and mechanistic understanding. While deep learning has advanced computational DTI prediction, most existing methods rely primarily on molecular structural representations, including drug structures and protein sequences, while overlooking cellular phenotypes that reflect downstream biological effects. Cell Painting enables high-content morphological profiling that captures systems-level responses to chemical and genetic perturbations but remains underutilized in DTI modeling. Integrating molecular information with cellular phenotypes offers an opportunity to improve both predictive performance and biological interpretability. Results We propose a two-stage contrastive learning framework integrating drug structures, protein sequences, and Cell Painting morphological profiles into a unified embedding space. Stage 1 learns modality-specific representations independently from structure-based and image-based data; Stage 2 aligns these via multi-positive contrastive learning to bridge molecular structural information with cellular phenotypes. Cross-modal retrieval achieves median Recall@10 values of 0.77 (random split) and 0.33 (scaffold split), outperforming bilinear and random baselines. In external DTI prediction on the BIOSNAP dataset, our model achieves an AUC of 0.92 with image-based representations and 0.90 under structure-only settings, surpassing existing methods. Model interpretation via integrated gradients reveals pathway-specific morphological signatures associated with drug targets, providing biologically interpretable insights into drug mechanisms. Availability https://github.com/YJRubyLai/Unified-DTI

Read PDF

Similar papers

Jul 2026

From Cellular Responses to Pharmacological Domains: Multimodal Zero-Shot Drug Representation Learning

PMRD separates mechanism-consistent factors from modality-specific information and constructs a consensus response domain across three modalities and combines complementary representations through reliability-aware multiview retrieval and supports PMRD as an effective framework for mechanism-aware multimodal drug representation learning.

Jintao Huang, Lu Leng, Ziyuan Yang · 0 citations
Preprint Aug 2026

Learning Molecular Representations from Cellular Phenotypes with Structure Preservation

PhenMol disentangles molecular and cellular representations into shared and private components, enabling phenotype-guided alignment while preserving chemical structures through a dedicated molecular branch, and improves molecular property prediction across 270 bioactivity tasks, molecule--phenotype retrieval, and clinical trial outcome prediction.

Xuan Lin, Jingyu Sheng, Tengfei Ma et al. · 0 citations
Review Jul 2026

Self-Supervised Learning for Molecular Property Prediction: Methods, Multimodal Insights, and Benchmark Comparisons

This review provides a systematic overview of recent advances in SSL-based molecular property prediction and analyzes how multimodal molecular representation learning by integrating sequence, graph, three-dimensional structure, and textual information can improve the quality and expressiveness of molecular representations.

Shuning Yang, Lei Deng · 0 citations
Aug 2026

Dual representation learning-based drug synergy prediction via sequence and molecular network reconstruction

The DRL-DSP is proposed, a novel dual representation learning framework designed to enhance drug synergy prediction by integrating molecular-level features from SMILES sequences with graph-level relational information from reconstructed molecular networks.

Juanzi Zhou, Xiaoliang Yang, Yin Zhang et al. · 0 citations
Jul 2026

Deep cross-modal representation fusion learning for enhanced drug target affinity prediction.

Prediction of Drug Target Affinity (DTA) is essential for accelerating computational drug discovery and reducing experimental costs. However, traditional experimental approaches for DTA estimation are resource-intensive and are further challenged by the structural flexibility of both drugs and target proteins. In this work, we propose the PCBERT-GAT-DFFNN-DTA model, a three-stage deep cross-modal representation fusion framework for accurate DTA prediction. In the first stage, variable-length protein sequences are transformed into contextual representations using ProtBERT to obtain fixed-size protein embeddings. Drug molecules are represented using two modalities: sequence-based embeddings generated from ChemBERT and structure-based embeddings learned from molecular graphs using a Graph Attention Network (GAT). In the second stage, each modality is processed through dedicated subnetworks to refine features and reduce dimensionality while preserving modality-specific information. In the final stage, the refined representations are fused and passed to a Deep Feed-Forward Neural Network (DFFNN) to predict drug target binding affinity. The proposed model consistently outperformed most baseline methods under the S1-S3 evaluation settings across the benchmark datasets. Under the more challenging S4 blind setting, the model achieved strong performance on the KIBA dataset and competitive results on the Davis and Metz datasets. Compared with LLMDTA, the proposed approach achieves significant improvements in R2 scores across all datasets, demonstrating its effectiveness in learning complex drug protein interactions for reliable DTA prediction.

Essmily Simon, Sanjay S. Bankapur · 0 citations
Open access Jul 2026

Momentum contrast-enhanced multimodal representation learning for drug synergy prediction

Abstract Motivation Accurate prediction of synergistic drug combinations can accelerate anticancer combination discovery. Existing methods inadequately model higher order drug–drug–cell-line interactions and drug–disease associations and remain sensitive to sparse and noisy multiomics data, limiting generalization to unseen cell lines and drug combinations. Results We present Momentum Contrast (MoCo)-MultiSynergy, a multimodal framework that combines modality-specific momentum contrastive learning with heterogeneous hypergraph modeling. The hypergraph represents synergistic drug–drug–cell-line triplets and drug–disease associations, while gated residual propagation refines node representations. MoCo modules regularize encoded drug and cell-line representations using latent feature masking and Gaussian perturbation. On the O’Neil and NCI-ALMANAC datasets, MoCo-MultiSynergy achieves the highest AUROC and AUPRC across the evaluated settings, with the largest gains when generalizing to unseen cell lines and drug combinations. Availability and implementation Source code is available at https://github.com/27167199/MoCo-MultiSynergy.

Yunxia Gu, Xindi Huang, Lifen Shi et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.