Skip to content
Open access

Flu Mutation Explorer: an Interactive Platform for Mapping Host Adaptation Mutations in Influenza A Viruses

Jul 2026 · bioRxiv · 0 citations · 24 references
Biology

TL;DR

The Flu Mutation Explorer is presented, an interactive web application that combines large-scale influenza phylogenies with a manually curated database of reported mammalian adaptation mutations, to enable the exploration and interpretation of IAV genetic variation.

Abstract

A rapid expansion of influenza A virus (IAV) genome sequencing has transformed global surveillance but has also created major challenges for interpreting the biological significance of viral mutations, particularly amino acid replacements associated with host adaptation. Resources have been created to support mutation annotation and phylogenetic analysis, but there is a need for a tool that integrates experimentally derived phenotypic evidence with evolutionary context in a framework suitable for users without prior training in bioinformatics. Here, we present the Flu Mutation Explorer, an interactive web application that combines large-scale influenza phylogenies with a manually curated database of reported mammalian adaptation mutations, to enable the exploration and interpretation of IAV genetic variation. The underlying database comprises over 1.5 million publicly available IAV sequences and over 1000 mutations associated with mammalian adaptation. The Flu Mutation Explorer enables users to query protein sequences, visualise amino acid distributions across viral lineages, examine host-specific conservation patterns, and identify adaptation mutation with links to supporting literature. We include case studies which demonstrate the platform’s use in assessing amino acid conservation at sites of interest and in rapidly identifying candidate mammalian adaptation mutations during the ongoing H5N1 panzootic. By integrating genomic, phylogenetic, and functional information into an intuitive interface, the Flu Mutation Explorer lowers the barriers to interpreting influenza sequences for specialists and non-specialists alike.

Read PDF

Similar papers

Review Aug 2026

AI-enabled viral genomics: from virus discovery to host prediction and emerging variant forecasting.

The rapid expansion of metagenomic sequencing has generated vast repositories of viral sequence data that far outpace our capacity to interpret them using conventional approaches. Highly divergent sequences, sparse functional annotation, and taxonomically uneven sampling present fundamental challenges for reference-dependent methods, which lose sensitivity precisely for novel and understudied viruses with high public health relevance. Artificial intelligence (AI) provides a new avenue to address these challenges by enabling predictive inference from viral genomes and proteins while reducing dependence on sequence similarity. In this Review, we discuss representative advances in AI for virus discovery, taxonomic classification and functional annotation, prediction of host range and zoonotic potential, and efforts toward forecasting emerging variants. These advances are transforming viral genomics from a largely descriptive discipline into one with increasing predictive capability. We also critically assess the major challenges that constrain current approaches, including the availability of high-quality and representative datasets, rigorous model evaluation, biological interpretability and responsible governance for increasingly capable AI models.

Jincan Ke, Heng Rong, Yin Chen et al. · 0 citations
Open access Jul 2026

EscaPRRS-ORF5: a structure-aware evolutionary framework for prioritizing immune escape-prone variants in porcine reproductive and respiratory syndrome virus

Abstract Motivation Porcine Reproductive and Respiratory Syndrome Virus (PRRSV) is a rapidly evolving RNA virus causing significant economic losses, posing a formidable challenge to vaccine efficacy due to its high mutational variability and immune escape. As the viral mutants evolve, their ability to sustain in population is driven by a range of host biology factors such as receptor binding, fusion, and uncoating. Existing tools that predict viral fitness and escape propensities rely heavily on extensive, up-to-date sequence data and lack integration of biochemical host interactions, limiting mechanistic understanding of the mutational landscape. We introduce Esca, a sequence-only toolchain framework that identifies immune escape-prone residues by exhaustively scanning each residue position for all amino acid substitutions using a Bayesian Variational Autoencoder (VAE) trained on protein language model embeddings. We demonstrate Esca on the GP5(ORF5) glycoprotein of PRRSV (EscaPRRS-ORF5) by training on ESM-2 embeddings of 32 146 GP5 sequences (2015–2022) spanning 140 sub-lineages. Results Despite being trained only on GP5 sequence data, EscaPRRS-ORF5 recovered 85.7% of the surface-exposed receptor binding interfaces as escape-prone regions. We use a mutation-sensitive fitness scoring scheme that goes beyond Hamming distances, to predict antibody escape tendencies, supporting surveillance of (re) emerging PRRSV variants. We do not claim that ORF5 alone captures PRRSV evolution or serves as a surveillance endpoint; rather, Esca offers a scalable path toward whole-genome, structure-aware surveillance. Availability and implementation EscaPRRS-ORF5 is freely available at https://doi.org/10.6084/m9.figshare.32661033 with an interactive Colab notebook at https://colab.research.google.com/drive/1TEgzAhPwvNAZ01VXeJbIFibfri2jnDA5? usp=sharing.

Ratul Chowdhury, Vaishnavey, Supantha Dey et al. · 0 citations
#protein folding Open access Aug 2026

Tree-aware conditional language modeling recovers mutational patterns of viral evolution

EvoPLM-Tree, a tree-aware conditional autoregressive language model that predicts descendant protein sequences from ancestral sequences together with phylogenetically derived evolutionary features, provides a framework for modeling protein evolution along phylogenetic lineages and prioritizing plausible future mutations from genomic surveillance data.

Polina V. Polunina, Wolfgang Maier, Alan F. Rubin · 0 citations
Open access Sep 2026

NGS-RFA/FRA: A High Throughput Experimental and Computational Pipeline for Selective Mutational Scanning in Parallel

Influenza viruses evade vaccine and infection mediated immunity by accumulating mutations in their hemagglutinin (HA) protein. Predicting this evolution might be possible via selective mutational scanning (SMS) – the generation of many specific mutants of interest from currently circulating viruses and characterizing their escape potential and fitness with high accuracy (Mögling 2016). However, this task is challenging, even when focusing on a reduced set of key HA positions (Koel et al. 2013). Here we describe a high-throughput SMS method to address this challenge. Our approach consists of a three-stage pipeline: (1) a parallel optimized virus rescue process that generates balanced target mutant virus libraries (2) an assay to assess replicative fitness and neutralisation of these variants as a mixture, and (3) a bespoke statistical model to quantify statistically significant differences between these observables. We tested the pipeline on libraries of up to 134 variants finding excellent correlation to classical hemagglutination inhibition (HI) and plaque growth assays used to assess antigenic phenotype and replicative fitness respectively, as well as remarkable repeatability overall. Notably, the method reduces the timeline required to carry out such assessments with classical methods from about a year to several weeks. By enabling rapid and efficient characterization of influenza virus variants, this approach has the potential to greatly enhance surveillance efforts, transforming reactive monitoring into proactive forecasting.

Sina Tureli, T. Bestebroer, S. James et al. · 0 citations
Open access Aug 2026

Anniemap: Vector Search for Viral Short Read Alignment

Background The process of aligning sequencing reads to a reference genome is a foundational step in genomic analysis, underpinning tasks from variant detection to pathogen surveillance. In viral genomics, however, this problem becomes substantially more challenging: viral sequences are often present at low abundance within host-dominated samples and can differ markedly from available references due to rapid mutation and population heterogeneity. These characteristics reduce the effectiveness of conventional seed-and-extend aligners, which typically rely on long exact or near-exact matches to anchor alignments. Even modest sequence divergence or sequencing errors can disrupt such seeds, particularly for short reads, leading to missed alignments. The central challenge in this setting is maintaining robust alignment under high divergence without sacrificing efficiency. Results We introduce Anniemap, a vector search–based approach to viral short-read sequence alignment. Anniemap represents reads and reference sequences as binary vectors and performs approximate nearest-neighbour search using Facebook AI Similarity Search (FAISS) to efficiently identify candidate mappings. Anniemap was compared with the well-established alignment tools Bowtie2 and BWA-MEM2 across a diverse set of viral genomes and read lengths using both simulated and real sequencing data. Anniemap achieved higher sensitivity and throughput in almost all evaluated scenarios, with the most substantial improvements in sensitivity observed for highly divergent genomes, such as Hepatitis C virus (HCV) and Human Immunodeficiency Virus (HIV). Conclusions By measuring vector similarity rather than relying on long exact seed matches, Anniemap provides greater robustness to sequencing errors and genomic mutations. This property is particularly advantageous for viral genomes, where substantial sequence divergence is common. Further work is required to efficiently extend vector-based search for read alignment beyond viral genomes.

D. J. van Zyl, H. Tegally, C. Baxter et al. · 0 citations
Open access Aug 2026

Alignment-free prediction of cross-reactivity in influenza A (H3N2) anticipates antigenic drift

Since its introduction in 1968, Influenza A (H3N2) has undergone continuous antigenic evolution, necessitating frequent vaccine updates. To predict antigenicity and characterize antigenic drift without multiple sequence alignments, we present FluEmbed, a computational framework that leverages protein language models. FluEmbed accurately quantified the antigenic impact of viral evolution from RNA sequences, achieving strong predictive performance against hemagglutination inhibition (HI) assay titers (Spearman correlation: ρ = 0.67–0.80). FluEmbed also outperformed sequence-distance baselines (e.g., Hamming and BLOSUM62) and phylogenetic tree-based models that require sequence alignment. Using this model, we conducted in-silico mutagenesis experiments to identify site/amino acid combinations that differentially impacted antigenicity. To systematically investigate how specific mutations influence immune escape, we defined two classes of mutations: ‘constrained’, where only the most likely amino acid changes at historically mutation-prone sites were considered (thereby limiting the mutation space) and ‘unconstrained’, where all possible substitutions were allowed, providing a full exploration of potential antigenic shifts. Constrained mutations often confer limited antigenic changes, whereas unconstrained mutations exhibit greater escape potential, particularly outside the dominant viral lineages. Notably, 3C.2a was the only major lineage in which constrained and unconstrained mutations showed no significant difference (p ≈ 0.95), suggesting ongoing intra-clade competition rather than inter-lineage antigenic replacement. By enabling rapid, alignment-free antigenic prediction directly from sequence data, FluEmbed could complement traditional HI assays in real-time influenza surveillance and inform vaccine strain selection decisions.

A. Forna, Lambodhar Damodaran, C. Gunning et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.