Skip to content
Review Open access

From sites to structure to serology: a roadmap for structure-aware molecular evolution of antigenically evolving viruses

Jul 2026 · Journal of Virology · Vol 100 · 0 citations · 202 references
Medicine

TL;DR

A practical framework linking sites, structure, and serology for viruses in which antigenic evolution is a major component of immune escape and lineage turnover is outlined and how integrating genomic surveillance data, phylogenetics, structural analysis, and predictive modeling could support more prospective variant assessment and improved vaccine and therapeutic design is discussed.

Abstract

ABSTRACT The genomic deluge has pushed viral molecular evolution into a site-resolved era. For antigenically evolving viruses such as influenza and severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2), dense genomic sampling now supports mutation-annotated phylogenies and per-site estimates of mutation and substitution processes. These data highlight strong effects of sequence context, genomic region, RNA structure, and protein-level constraints that are blurred by classic uniform substitution models. In parallel, accurate structure prediction and emerging structure-aware phylogenetic and machine-learning approaches provide practical ways to map mutations onto three-dimensional constraints, identify structurally plausible escape routes, and interpret evolutionary rate variation through solvent exposure, packing, stability, glycosylation, receptor-binding interfaces, and epitope geometry. Finally, antigenic cartography translates some forms of genetic change into an epidemiologically meaningful phenotype—antigenic distance—while predictive modeling increasingly enables sequence-to-antigenicity inference for variants that have not yet been tested experimentally. Here, we outline a practical framework linking sites, structure, and serology for viruses in which antigenic evolution is a major component of immune escape and lineage turnover; highlight why genetic and antigenic “clocks” can diverge; and discuss how integrating genomic surveillance data, phylogenetics, structural analysis, and predictive modeling could support more prospective variant assessment and improved vaccine and therapeutic design.

Read PDF

Similar papers

Open access Aug 2026

Alignment-free prediction of cross-reactivity in influenza A (H3N2) anticipates antigenic drift

Since its introduction in 1968, Influenza A (H3N2) has undergone continuous antigenic evolution, necessitating frequent vaccine updates. To predict antigenicity and characterize antigenic drift without multiple sequence alignments, we present FluEmbed, a computational framework that leverages protein language models. FluEmbed accurately quantified the antigenic impact of viral evolution from RNA sequences, achieving strong predictive performance against hemagglutination inhibition (HI) assay titers (Spearman correlation: ρ = 0.67–0.80). FluEmbed also outperformed sequence-distance baselines (e.g., Hamming and BLOSUM62) and phylogenetic tree-based models that require sequence alignment. Using this model, we conducted in-silico mutagenesis experiments to identify site/amino acid combinations that differentially impacted antigenicity. To systematically investigate how specific mutations influence immune escape, we defined two classes of mutations: ‘constrained’, where only the most likely amino acid changes at historically mutation-prone sites were considered (thereby limiting the mutation space) and ‘unconstrained’, where all possible substitutions were allowed, providing a full exploration of potential antigenic shifts. Constrained mutations often confer limited antigenic changes, whereas unconstrained mutations exhibit greater escape potential, particularly outside the dominant viral lineages. Notably, 3C.2a was the only major lineage in which constrained and unconstrained mutations showed no significant difference (p ≈ 0.95), suggesting ongoing intra-clade competition rather than inter-lineage antigenic replacement. By enabling rapid, alignment-free antigenic prediction directly from sequence data, FluEmbed could complement traditional HI assays in real-time influenza surveillance and inform vaccine strain selection decisions.

A. Forna, Lambodhar Damodaran, C. Gunning et al. · 1 citation
#protein folding Open access Aug 2026

Tree-aware conditional language modeling recovers mutational patterns of viral evolution

EvoPLM-Tree, a tree-aware conditional autoregressive language model that predicts descendant protein sequences from ancestral sequences together with phylogenetically derived evolutionary features, provides a framework for modeling protein evolution along phylogenetic lineages and prioritizing plausible future mutations from genomic surveillance data.

Polina V. Polunina, Wolfgang Maier, Alan F. Rubin · 0 citations
Open access Aug 2026

Conserved influenza A epitope candidate regions and a benchmark of ESM-2 sequence features

Influenza A virus antigenic drift forces annual vaccine reformulation, motivating the search for conserved epitope candidates that could support broadly protective vaccines. We systematically screened influenza A virus sequences (H1N1, H3N2, H5N1; nine viral proteins) to define 98 conserved candidate regions, 38 of which were identical across the H1N1, H3N2, and H5N1 consensus sequences — all in the polymerase complex and nucleoprotein (PB2, PB1, PA, NP) — whereas the ten surface-glycoprotein (HA/NA) candidates were subtype-specific. We then benchmarked two protein-language-model (ESM-2) features against alignment conservation. Group-masked log-probability correlated moderately with MSA conservation (Spearman ρ = 0.25–0.39 for HA) but provided no incremental value for T-cell epitope discrimination (ΔAUROC +0.004, p = 0.46); attention-derived contact-density was not a valid solvent-accessibility proxy. A curated antibody-epitope benchmark (22 clusters, 5 neutralization-supported) was underpowered for a high-confidence B-cell test. We document data-quality and reproducibility pitfalls (length heterogeneity, coordinate mapping, and pseudoreplication) and release the auditable benchmark. These results provide an auditable candidate resource and show that, in the evaluated benchmarks, ESM-2 sequence scores did not improve epitope prioritization beyond alignment-derived conservation.

Qingxiu Li, Zhen-Jun Li · 0 citations
Open access Jul 2026

Beyond Invariable Sites: Using Evolutionary Stasis to Map Multilayered Constraints on the Evolution of Viral and Mammalian Genomes

Abstract The quantification of genomic conservation has progressed from foundational statistical modeling of evolutionary rates to state-of-the-art deep learning architectures. However, a major resolution gap remains at the zero-rate origin, where standard selection inference tools fail to distinguish between sites that are invariant due to chance (stochastic invariance) or low substitution opportunity and those that are invariant due to extreme purifying selection. We present B-STILL (Bayesian Significance Test of Invariant Low Likelihoods), a hierarchical Bayesian framework designed to resolve the selective landscape of protein-coding genes near the zero-rate limit. By leveraging gene-level rate distributions (prior calibration) and modeling codon-site-specific substitution opportunities (determined by genetic-code degeneracy and nucleotide substitution biases), B-STILL quantifies the statistical significance of observed stasis. We define a rate-based stasis threshold to identify evolutionary stasis anchors (ESAs)—sites where the upper bound on the evolutionary rate is statistically constrained relative to the background rate of the gene due to extreme purifying selection. Validation against clinical and pathogen datasets confirms that ESAs are strong predictors of biological fitness and pathogenicity. Applying B-STILL across viral and mammalian genomes, we identify thousands of significantly clustered ESAs that map to known functional domains and uncharacterized structural motifs. These results establish B-STILL as a scalable, statistically rigorous framework for high-resolution genomic annotation, converting previously uninformative invariant sites into precise markers of extreme evolutionary constraint.

Sergei L. Kosakovsky Pond, Hannah Verdonk, Steven Weaver et al. · 0 citations
Open access Sep 2026

Forecasting viral evolution from phylogenetic trees

Viral mutation forecasting plays a key role in pandemic preparedness by enabling researchers to anticipate novel variants and design proactive interventions. Evolutionary histories, represented as phylogenetic trees, offer key insights into the emergence of past and present strains, yet their role in predicting future sequence changes remains largely unexplored. We introduce antiGen, a machine learning model that forecasts the evolutionary future of viruses by learning from their evolutionary past. antiGen achieves state-of-the-art performance for predicting mutations to the SARS-CoV-2 spike protein, anticipating never-before-seen mutations and mutations that emerge years after the model’s training window. antiGen-forecasted spike mutations also retain pseudoviral infectivity in vitro. Moreover, antiGen demonstrates leading predictive performance on surface proteins of influenza virus, respiratory syncytial virus, and dengue virus despite far less available sequencing data. Viral evolution models that explicitly learn from phylogenetic structure offer a valuable resource for applications ranging from epidemiological modeling to therapeutic development.

I. Specht, Soyoon Park, Seyone Chithrananda et al. · 0 citations
Open access Jul 2026

Epistasis facilitates the long-term antigenic evolution of the influenza B virus hemagglutinin

The antigenic drift of viral glycoproteins must be balanced by purifying selection pressure to maintain functionality. Understanding these evolutionary processes is key to predicting and combating viral evolution but is primarily based on influenza A(H3N2), which may limit generalisability. By characterising the influenza B virus haemagglutinin (HA) over 8 decades of circulation in humans, we found continuous genetic diversification, punctuated with antigenic changes that did not follow a linear path in antigenic space. Antigenic change is primarily underpinned by re-occurring mutations and deletions at positions 136, 150, 162-165, 197 and 203. These residues form complex epistatic networks that modulate the antigenic impact of mutation recycling. They also generate permissive backbones on which immune escape can emerge with limited replicative fitness cost. Our study identifies critical similarities and differences with A(H3N2) evolution and demonstrates the role of epistasis in balancing antigenic novelty with viral fitness. Our findings and genetic, antigenic and phenotypic datasets support the development of genotype-to-phenotype prediction tools, but such predictions need to capture the complex outcomes of epistasis.

Lara S. U. Schwab, Ruo-Peng Xie, Ellie Reilly et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.