T cell receptor (TCR) binding to peptides presented by major histocompatibility complex (MHC) molecules is a key step in T cell activation, and forms the basis of adaptive immunity. Predicting this specificity is therefore essential to developing effective TCR-based immunotherapies and vaccines. Despite its clinical relevance, predicting TCR-pMHC specificity for previously unseen peptides remains an open problem, with structural modeling so far the only strategy showing any predictive power in this setting. In this study, we find that this limited performance is substantially driven by label noise in the data used to train and evaluate these methods, an effect that has so far been largely underexplored. Using an AlphaFold3-based pipeline adapted for TCR-pMHC structural modeling, we achieve state-of-the-art specificity prediction, outperforming AlphaFold2.3-based and sequence based methods, and performing at par with the leading Immrep2025 competition submission. Combining this pipeline with a cluster-based denoising algorithm, we show that removing mislabeled points from a large specificity dataset increased binder ranking accuracy by more than 70% relative to the full dataset. Together, these results highlight label noise as a major factor limiting the performance that any method in this field can achieve, and show that combining structural modeling with label denoising substantially improves TCR-pMHC specificity prediction, making such approaches an attractive complement to current sequence-based approaches for refining TCR target selection.
Pilar Ballesteros-Cuartero, J. Lund, Morten Nielsen· bioRxiv· 1 citation· ⚡1
Cancer epitopes, the molecular structures recognized by T and B cells at the tumor interface, are central to understanding antitumor immunity and developing immunotherapies. Yet despite the rapid growth of cancer immunology data, a comprehensive, continuously updated, and accessible resource for cancer epitope data has been lacking. The Cancer Epitope Database and Analysis Resource (CEDAR, cedar.iedb.org) was established in 2021 to fill this gap, providing curated experimental epitope data alongside a suite of cancer-specific computational tools for epitope prediction and analysis. Built on the validated infrastructure of the Immune Epitope Database (IEDB), CEDAR integrates cancer epitope data with biological, immunological, and clinical context, enabling researchers to explore immune recognition of tumors, identify candidate targets for immunotherapy, and benchmark prediction methods. Here we describe CEDAR’s current capabilities, report on progress in curation, database development, and tool availability, and outline the opportunities and challenges ahead for expanding its scope and utility to the cancer research community.
Zeynep Koşaloğlu-Yalçın, Ibel Carri, Daniel Marrama et al.· Frontiers in Oncology· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.