Interpretable biophysical neural networks of transcriptional activation domains separate roles of protein abundance and coactivator binding.
Abstract
Deep neural networks have improved many difficult prediction tasks in biology, but it remains challenging to interpret these networks and learn the molecular mechanisms. Here, we address interpretation challenges by building biophysical neural networks for predicting activation domains, the regions within transcription factors (TFs) that recruit coactivators to drive gene expression. Deep neural networks can now accurately predict acidic activation domains from protein sequences, but these predictors are difficult to interpret. We designed shallow neural networks that incorporated biophysical models and visualized the parameters directly. We found two ways that the arrangement of residues (i.e., sequence grammar) controls function: (1) C-terminal hydrophobic residues increase coactivator binding and decrease protein abundance, and (2) acidic residues at the N terminus promote coactivator binding, while acidic residues at the C terminus promote TF abundance. We demonstrate how combining biophysical and deep neural networks maximizes prediction accuracy and interpretability, revealing biological mechanisms across datasets. A record of this paper's transparent peer review process is included in the supplemental information.