A machine learning-based approach to evolve new anti-PD-1 antibodies within a chemically informed latent space using a conditional kernel-elastic autoencoder (CKEA) between nivolumab and pembrolizumab to preserve favorable therapeutic features while exploring variants with different potency.
Abstract
The human immune system excels at generating highly effective antibodies through natural selection and somatic hypermutation, but adapting these antibodies for therapeutic use, referred to as "antibody medicine-likeness", requires careful consideration of biochemical and physiological properties. Traditional redesign methods are often slow and limited in scope. In this study, we introduce a machine learning-based approach to evolve new anti-PD-1 antibodies within a chemically informed latent space using a conditional kernel-elastic autoencoder (CKEA) between nivolumab and pembrolizumab, both of which bind the FG-loop "hotspot" of PD-1 in the most distantly related orientations, differing by 174°. This generative framework is designed to preserve favorable therapeutic features while exploring variants with different potency, ultimately for improved potency. To evaluate structural and functional viability, we performed molecular dynamics (MD) simulations of the generated antibody – PD-1 complexes and described their MD properties. These simulations reveal detailed free-energy landscapes and identify stable binding conformations, providing a strong basis for experimental validation. To validate our designs, we expressed and experimentally tested the antibodies for binding affinity to PD-1. Upon expression and purification, three out of six designed antibodies exhibited some binding to PD-1, whose properties could likely be improved using other computational saturation mutagenesis or laboratory evolution. Our results demonstrate the potential of artificial intelligence (AI)-guided interpolation methods to generate novel, high-affinity antibodies with therapeutic promise, offering a powerful strategy for next-generation antibody development. SYNOPSIS TOC Combination of MD simulations with machine learning algorithms could revolutionize the antibody-breeding sciences to lead to new antibody discovery that is compatible with or better than naturally occurring antibodies. Topics of Content (TOC) Fingerprints of R86 finger of PD-1 are recognized by our designed P2N_2 anti-PD-1 antibody according to our MD simulations
CLDN18.2 is a promising tumor-specific antigen; however, the development of therapeutic antibodies against it is challenged by the need for simultaneous optimization of affinity and developability. To address this, we present cdrGPT, a deep learning framework based on GPT-2 for de novo generation of complementarity-determining region H3 (CDRH3) sequences. Our approach integrates pre-training on the Observed Antibody Space (OAS) database with structural templating derived from the known antibody zolbetuximab. Generated sequences were iteratively refined through rejection sampling and fine-tuned against a multi-parameter objective function encompassing predicted affinity and MHC class II binding risk. From an initial set of 50,000 sequences, this screening pipeline yielded 313 high-confidence candidates. Subsequent analysis using evolutionary scale modeling 2 (ESM2) embeddings, principal component analysis (PCA), and clustering revealed three structurally distinct clusters, with intra-cluster cosine similarities exceeding 0.99. Validation of seven representative sequences from the dominant cluster using AlphaFold3 confirmed high structural fidelity to the zolbetuximab template, demonstrating a root mean square deviation (RMSD) of 1.331 Å for the CDRH3 loop and positional deviations of less than 0.4 Å for key paratope residues. These results indicate that the designed variants preserve the core binding mode of the parent antibody. This study establishes a feasible pipeline for integrating AI-generated CDRH3 loops into functional antibody scaffolds, providing a foundation for the accelerated development of therapeutics targeting CLDN18.2 and other clinically relevant antigens.
Tao Qu, Lingyan Yuan, Wei-Ran Cui et al.· PLoS Computational Biology· 0 citations
This work designs a biological prior-guided feature fusion framework that integrates pseudo-structural epitope knowledge and CDR-specific attention mechanisms via a mixture-of-experts architecture to effectively capture complex binding landscapes in antibody screening and drug residence time analysis.
G. Luo, Junkai Wang, Sizhe Zhang et al.· Bioinformatics· 0 citations
The ImmunoFoundation Model (IFM), a multimodal deep learning system that integrates not only peptide sequences, 3D molecular structures, and biochemical properties but also TCR-MHC-peptide to achieve superior immunogenicity prediction and enable peptide optimization for therapeutic applications is developed.
Smita Krishnaswamy, J. Rocha, Hiren Madhu et al.· Journal of Immunology· 0 citations
Antibodies raised against human targets often fail to recognize their animal orthologs, limiting preclinical evaluation in relevant models. We developed a Deep Mutational Scanning (DMS)-coupled deep learning strategy to engineer potent cross-reactive antibodies with minimal sequence divergence. Starting from C4, a fully human anti-PD-L1 antibody with weak recognition of murine PD-L1, DMS identified substitutions that improved binding to both human and mouse antigens. Conventional recombination of beneficial mutations generated highly cross-reactive antibodies but required 13 to 15 substitutions. To reduce this mutational burden, a deep learning model trained on DMS-derived sequence-binding data was used to identify minimal mutation combinations predicted to retain high affinity. This approach yielded variants carrying only 4 to 5 substitutions, with in vitro and cellular binding properties comparable to highly mutated antibodies. Epitope mapping, structural modeling and in vivo assessment further confirmed that these engineered antibodies retained PD-1/PD-L1 blockade and demonstrated therapeutic activity in a mouse tumor model.
H. Dorison, Anne-Laure Grindel, François Thenier et al.· bioRxiv· 0 citations
Personalized neoantigen prediction is challenging due to the scarcity of positive samples, the noise of the experimental data, the severe class imbalance trait and the complex of immunogenicity features. Prior arts, such as linear regression and XGBoost fail to model long-range dependencies and contextual relationships within peptide features, therefore the performance of neoantigen positive recall rate is limited. In this paper, we present a novel deep learning framework based on Transformer, coined as TransNRank. By leveraging the self-attention mechanism, our model captures both local and global feature contexts, enabling more accurate recognition of immunogenic neoantigens. A positive-aware training objective is utilized to handle the class imbalance problem, assigning more weights to those few positive samples. Extensive experiments are performed on NCI, TESLA and HiTIDE datasets. Notably, our TransNRank can push the upper bound top 20 recall rate of neoantigen prediction from 46.9% (45 from 96) to 53.1% (51 from 96), while reducing the training epochs from 200 epochs to 20 epochs. Furthermore, we analyze the features contribution based on TransNRank and find that the mutation at anchor and TCGA expression level play an unexpected important role in neoantigen prediction, and removing insignificant features to reduce the input dimensionality of peptides does not drastically impair the overall performance of the model. Our paradigm not only streamlines the prediction pipeline but also sets a new state-of-the-art for neoantigen discovery, with broad implications for accurate immuno-oncology.
Zhiyin An, Yue-Nan Hou, Shumeng Duan et al.· 0 citations
Somatic hypermutation (SHM) involves activation-induced cytidine deaminase (AID)-mediated DNA targeting, followed by error-prone repair which introduces mutations. However, existing computational models do not reflect this, limiting their ability to model the mechanism of affinity maturation in silico. Here, we develop a biologically grounded, transformer-based framework that mirrors SHM. Our framework employs two independent antibody language models (AbLMs): an AID-like targeting model to select mutation sites and a DNA repair-like substitution model to predict resulting amino acids.
We sequenced memory B cell repertoires from eight healthy donors using high-accuracy bulk NGS of unpaired VH/VL chains. These data were used for pretraining the two AbLMs to perform in silico SHM. Using AntiBERTa2-derived metrics as a correlate for native-like antibodies, we compared model-generated sequences to the natural human immune repertoire.
Bulk NGS produced ∼84 million high-quality, productive memory B cell receptor sequences for AbLM training. Model predicted mutations matched the spatial distribution and overall load of true SHM. AntiBERTa2 embedding projections and log likelihood scoring revealed that model-generated sequences closely resemble native antibodies and strongly mimic SHM mutational patterns. Model-generated sequences also exhibited significantly increased levels of expression in HEK293F cells, indicating learning of beneficial mutations that enhance protein fitness.
Our framework models SHM to generate native-like human antibodies, providing a biologically grounded tool for guiding design within the natural sequence space. Unlike motif-based or single-stage SHM simulators, our framework explicitly decouples targeting and substitution to enable faithful reproduction of both hotspot localization and amino acid substitution. This approach advances in silico modeling of humoral immunity and may reveal new mechanistic insights into antibody clonal dynamics.
Endowed Fellowship in the Skaggs Graduate School of Chemical and Biological Sciences, Achievement Rewards for College Scientists (ARCS) Foundation San Diego Chapter, National Institute of Allergy and Infectious Diseases (NIAID)
Computational and Systems Immunology (COMP)
Karenna Ng, B. Némoz, Bryan S. Briney· Journal of Immunology· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.