Skip to content

Author

Sung-Gwon Lee

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access Aug 2026

Interpretable prediction of nucleic acid-binding proteins using a protein language model

Abstract Motivation Accurate identification of DNA-binding proteins (DBPs) and RNA-binding proteins (RBPs) is critical for elucidating transcriptional and post-transcriptional regulatory mechanisms. However, existing computational approaches often rely on inferred labels or domain-specific annotations, which limit the subsequent generalizability. Results This study aimed to introduce transformer-based classifiers for human DBPs and RBPs that rely solely on protein sequence information without engineered features or domain constraints. The models were implemented using ESM-2 with low-rank adaptation (LoRA) fine-tuning and trained on experimentally validated datasets, including chromatin immunoprecipitation sequencing (ChIP-seq) annotations for DBPs and eCLIP annotations for RBPs. Next, to evaluate biological relevance, we computed value-aware attention (VAT) scores aggregated across transformer layers to interpret model focus. In 20-fold cross-validation, the DBP model achieved an area under the receiver operating characteristic curve (AUROC) of 0.84 with a Matthews correlation coefficient (MCC) of 0.40, while the RBP model achieved an AUROC of 0.92 with an MCC of 0.46. Proteins predicted as nucleic acid-binding were enriched for known binding domains, and inspection of attention distributions revealed preferential focus on annotated functional regions rather than non-binding segments. These results demonstrate that attention-based protein language models can accurately identify nucleic acid-binding proteins directly from sequence data. Moreover, these models reveal biologically meaningful sequence determinants of binding, establishing an interpretable and scalable framework for proteome-wide characterization of protein–nucleic acid interactions. Availability and implementation Code is available on GitHub (https://github.com/CSB-hub/DRBP).

Hanjin Kim, Sung-Gwon Lee, Joo-Seong Oh et al. · 0 citations
Open access Jul 2026

An openly licensed benchmark and per-gene calibration map for missense pathogenicity predictors on activating cancer drivers

Missense pathogenicity predictors such as AlphaMissense are increasingly used in clinical variant interpretation, yet they are trained on germline labels dominated by loss-of-function (LOF) variants. Using an openly licensed, reproducible benchmark of 768 Cancer Gene Census genes scored with 49 predictors (labels from CIViC, COSMIC, cancerhotspots, ClinVar and gnomAD), we show that 42 of 49 tools (86%) score oncogene, gain-of-function (GOF) variants worse than tumour-suppressor variants. This under-scoring is mechanistically characterized: missed drivers occupy low-conservation, solvent-exposed, non-destabilizing positions (phyloP 2.51 versus 7.89; relative solvent accessibility 0.671 versus 0.185; gene-clustered p = 4.8×10⁻²⁰ and 2.3×10⁻³⁵), and, counter-intuitively, the unsupervised and protein-language models now entering clinical use are the most affected. Per-gene oncogenic thresholds span 0.07–0.99, so a single global cut-off is mis-calibrated for most genes; we provide a per-gene calibration map. A cancer-calibrated stack (OncoCal) modestly improves discrimination over the best single tool (AUROC ≈ 0.93 versus 0.87), rescues drivers such as JAK2 V617F (0.334→0.57), and generalizes to independent deep mutational scanning data. We provide an openly licensed framework to interpret and recalibrate these tools in the somatic setting rather than a replacement predictor.

Sung-Gwon Lee · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.