Unsupervised detection of antimicrobial-resistance determinants by coupling protein-language-models and evolutionary signatures
Abstract
Antimicrobial resistance (AMR) is among the most pressing threats to global health, yet our ability to find resistance determinants is largely confined to what reference databases already contain: homology search and supervised classifiers recognize variants of known genes but are, by construction, blind to the larger environmental and clinical reservoir of determinants that have not yet been catalogued. To address this critical limitation, we tested whether resistance determinants can be flagged without using any resistance label, by leveraging evolutionary signatures that acquisition and adaptation leave in bacterial genomes. We describe a label-free, multi-view framework that scores every gene family of a pangenome on five orthogonal axes: protein-language-model novelty relative to known protein space, mobility/compositional anomaly, episodic positive selection, presence/absence homoplasy, and reconciliation-inferred horizontal transfer. These views were then combined based on a conjunctive (weighted geometric-mean) rule, so that only families implicated by several independent lines of evolutionary evidence score high. On a controlled simulation the conjunction recovers all planted determinants where no single view is specific. Applied without retraining to the Escherichia coli (n=150) and Klebsiella pneumoniae (n=150) pangenomes, the results confirm that known determinants are almost entirely accessory and concentrates them near the top of the ranking for K. pneumoniae (6.6-fold enrichment in the top 1%), but not for E. coli. Ablation shows presence/absence homoplasy carries most of the signal, that episodic selection is counterproductive, and that an equal-weighted conjunction is suboptimal. The framework offers a reproducible, database-independent shortlist of candidate determinants and a honest accounting of where evolutionary signal is, but is not sufficient on its own. Author summary Bacteria become resistant to antibiotics in ways we have not finished cataloguing, but the standard computational tools for finding resistance genes can only recognize genes that resemble ones already in a database. This makes genuinely new determinants, exactly the ones surveillance most needs to catch, the hardest to find. Here we evaluate an alternative strategy, based on the idea that resistance genes tend to leave specific evolutionary footprints: they are often acquired horizontally and jump between unrelated genomes, they sit on mobile elements, the same change arises repeatedly under drug pressure, and the proteins can look unusual to a protein-language model trained on natural sequences. We score every gene in a species’ pangenome on these signatures without ever telling the method which genes are resistance genes, and we keep only the genes that several independent signatures agree on. On simulated data, this strategy cleanly recovers planted resistance genes. However, on real Escherichia coli and Klebsiella pneumoniae genomes, it works partially: resistance genes are clearly enriched for known determinants in Klebsiella, but not in E. coli. Our results pinpoints which signatures help and which mislead, and highlight a transparent, database-free way to shortlist candidate resistance genes, together with a map of its current limitations, which in turn can be used to improve the approach.