Skip to content

Author

K. Dufault-Thompson

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access Aug 2026

Ancestral proteins trace the emergence of substrate specificity and oligomerization within bacterial DEDDy dinucleases

Nucleases are crucial for various bacterial processes, including genome maintenance and host defense. Deoxydinucleases (diDNases), a class of Gram-positive bacteria–specific nucleases associated with mobile genetic elements, are homologous to nanoRNase C (NrnC) in Gram-negative bacteria but exhibit notable differences: diDNases form dimers and cleave DNA dinucleotides, whereas NrnC forms octamers that process both RNA and DNA dinucleotides. The mechanism by which substrate specificity emerged, and whether it is linked to oligomerization, remained unknown. Here, we reconstructed a common ancestor of diDNases and NrnC orthologs that forms a dimer with intermediate preference for DNA. Structures of ancestral and extant dinucleases reveal gradual changes in conformation that gave rise to substrate preference, oligomeric state, and catalytic efficiency. These findings highlight how subtle, concerted structural modifications enable large-scale changes in molecular assembly and functional specialization, harnessing a conserved protein fold. DNA dinucleotide preference in the early ancestor and preservation of DNase activity in all extant enzymes strongly argue for a biological function of DNA dinucleotides.

Sofia Mortensen, Audrey A. Burnim, K. Dufault-Thompson et al. · 0 citations
Open access Sep 2026

LAMBDA: a prophage detection benchmark for genomic language models.

Transformer-based genomic sequence models represent an emerging frontier in computational biology. Yet, their embeddings have not yet shown the same level of predictive power as natural and protein language models, highlighting a gap between current implementations and theoretical promise. Existing benchmarks for DNA language models primarily focus on classifying regulatory elements in eukaryotic genomes, leaving open the fundamental question of whether these models learn sequence-level features across whole genomes. We introduce LAMBDA, a benchmark designed to rigorously evaluate genome language model embeddings through phage-bacteria sequence discrimination across four categories of increasing complexity: probing tasks, fine-tuning assessments, diagnostic tests, and genome-wide prophage detection. Our comprehensive analysis of current genomic language models provides insight into the importance of training data selection relative to model size, the need for domain-specific training, and the capabilities and limitations of genomic language models for detecting prophage sequences. This benchmark represents a challenging genomic annotation task in the bacterial domain and addresses a key computational problem with direct relevance to microbiology and medicine.

LeAnn M. Lindsey, Nicole L. Pershing, K. Dufault-Thompson et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.