Influenza A virus antigenic drift forces annual vaccine reformulation, motivating the search for conserved epitope candidates that could support broadly protective vaccines. We systematically screened influenza A virus sequences (H1N1, H3N2, H5N1; nine viral proteins) to define 98 conserved candidate regions, 38 of which were identical across the H1N1, H3N2, and H5N1 consensus sequences — all in the polymerase complex and nucleoprotein (PB2, PB1, PA, NP) — whereas the ten surface-glycoprotein (HA/NA) candidates were subtype-specific. We then benchmarked two protein-language-model (ESM-2) features against alignment conservation. Group-masked log-probability correlated moderately with MSA conservation (Spearman ρ = 0.25–0.39 for HA) but provided no incremental value for T-cell epitope discrimination (ΔAUROC +0.004, p = 0.46); attention-derived contact-density was not a valid solvent-accessibility proxy. A curated antibody-epitope benchmark (22 clusters, 5 neutralization-supported) was underpowered for a high-confidence B-cell test. We document data-quality and reproducibility pitfalls (length heterogeneity, coordinate mapping, and pseudoreplication) and release the auditable benchmark. These results provide an auditable candidate resource and show that, in the evaluated benchmarks, ESM-2 sequence scores did not improve epitope prioritization beyond alignment-derived conservation.
Small open reading frames (sORFs) are a potentially rich, yet error-prone, source of antimicrobial-peptide (AMP) candidates: short sequences are readily prioritized by AMP classifiers but may derive from incomplete gene calls. We developed a genome-context-aware discovery workflow that separates AMP-like sequence properties from evidence for a complete, recurrent coding locus. From 649,653 RefSeq assemblies representing 327 clinically relevant bacterial species, species-aware clustering and length filtering yielded 4,442,548 representative 10–100-aa sequences. AmpScanner v2, Macrel and AMPlify identified 585 non-haemolytic records supported by all three models. However, genome-context auditing of 11,918 mapped candidates showed that 529 of 536 mapped consensus candidates were supported exclusively by partial ORFs near contig termini. By contrast, 3,382 candidates had at least one complete non-edge occurrence; 1,069 recurred in ≥2 assemblies and 251 in ≥10 assemblies. We therefore assembled a 20-peptide panel through two explicitly labelled routes: sequence/structure-led selection (n=8) and genome-supported selection (n=12). Broth microdilution against Escherichia coli ATCC 25922 and Staphylococcus aureus ATCC 25923 identified low-micromolar activity in both routes. CAND_04141, a recurrent complete non-edge candidate, had the strongest combined profile (MICs of 4 and 2 μM, respectively), while CAND_07825 and CAND_04265 were also active at low micromolar concentrations. In plate-count MBC assays, all three advanced peptides achieved ≥3-log10 reductions at 128 μM. These findings show that high classifier agreement is not a substitute for genomic evidence and provide an auditable framework for prioritizing both synthetic AMP-like sequences and candidate genome-encoded peptides.
A sequence-traceable workflow linking proteome-scale eAMP discovery with structural prioritisation and experimental activity assessment is established, establishing a sequence-traceable workflow linking proteome-scale eAMP discovery with structural prioritisation and experimental activity assessment.
Qingxiu Li, Zhenjun Li· bioRxiv· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.