DeepPNI: a language- and graph-based model for mutation-driven protein–nucleic acid binding energetics
Abstract Protein–nucleic acid interactions (PNIs) are central to fundamental biological processes, and mutations can disrupt these interactions by altering local structural features and binding free energy. Here, we present DeepPNI, a deep learning regression model that integrates sequence- and structure-based features to estimate mutation-induced changes in binding free energy in protein–nucleic acid complexes. The model was developed using a comprehensive dataset of 1754 mutations spanning protein–DNA and protein–RNA complexes, representing one of the largest curated datasets for PNI binding free energy prediction. Structural features were encoded using an edge-aware relational graph convolutional network, while sequence features were represented using the Evolutionary Scale Modeling 2 protein language model. Despite the increased dataset size and heterogeneity, DeepPNI achieved an overall Pearson correlation coefficient of 0.76 in five-fold cross-validation. Consistent performance was observed across protein–DNA and protein–RNA subsets, datasets grouped by experimental temperature, and external blind test datasets, suggesting robustness against dataset heterogeneity. DeepPNI is freely available as a web server at https://research.iitbhilai.ac.in/molinfo/deeppni.