The results suggest that modest numbers of mutations suffice to reconstruct clonal tree topologies for typical numbers of clones, supporting subsampling as a general strategy for managing the challenges of ever-growing data.
Abstract
Abstract Motivation Phylogenetics faces a growing challenge from increasingly large and complicated data sets enabled by ever-improving sequencing technologies. The issue is particularly acute for somatic evolution studies, such as cancer cell lineages, where single-cell data sets may include tens of thousands of mutations in hundreds of thousands of genetically distinct cells. Simultaneously, the biological complexity of somatic evolution has led to complex phylogeny methods that struggle to scale to even modest data sizes. Results We explore the theoretical and empirical basis for one strategy for managing these large data sets: subsampling mutations for the computationally challenging phylogeny problem followed by faster placement of mutations on a putatively known guide tree. We specifically focus on determining the number of mutations sufficient to recover the true phylogeny at some level of resolution with high probability. We theoretically analyze variants of several common models that underlie popular tools for building clonal lineage trees. We further test these bounds through simulations of these models, extensions of them, and real biological datasets. The results suggest that modest numbers of mutations suffice to reconstruct clonal tree topologies for typical numbers of clones, supporting subsampling as a general strategy for managing the challenges of ever-growing data. Availability and implementation All analysis code and scripts used for data simulation are implemented in Python 3 and available at https://github.com/CMUSchwartzLab/mutation-subsampling.git.
This study compares two mutation models: one that incorporates recurrent mutations and another that allows only boundary mutations and shows that while tree topologies remain mostly unaffected, branch lengths and estimates of mutation bias and selection are substantially compromised.
Rui Borges, J. Hughes· Systematic Biology· 0 citations
Advances in genome sequencing have enabled phylogenomic studies involving tens or even hundreds of thousands of species. However, scalability remains a major computational challenge for statistically consistent species tree inference at this scale. ASTRAL, the most widely used coalescent-based species tree estimator, r...
Evolutionary biology has traditionally inferred process from patterns in extant organisms and the fossil record, leaving many foundational questions constrained by their historical nature. Over the past two decades, a growing family of synthetic approaches-high-throughput mapping such as deep mutational scanning, recon...
Xue-Ying C. Li, Paco Majic, C. Westmann· Genome Biology and Evolution· 0 citations
It is found that marine species tend to have lower ROH content, while fossorial species show significantly higher inbreeding levels, which suggests that habitat and foraging strata are much stronger predictors of IUCN status than genetics.
Many questions in population genetics require reconstructing evolutionary history through time, such as inferring how population structure has changed throughout the past. Yet, many existing approaches have only an implicit temporal component, using quantities such as allele frequency or haplotype length as rough proxi...
Yun Deng, Jonathan K. Pritchard, J. Spence· bioRxiv· 1 citation
Large phylogenetic trees are often unbalanced near the root, with a species-depauperate lineage sister to a far more species-rich clade. The species-poor branches, sometimes referred to as "early diverging," are often viewed as key sources of information about deep evolutionary history. As a result, these lineages occu...
Jeremy M. Beaulieu, Andrew J. Alverson, B. C. O'Meara· Systematic Biology· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.