The Linkage Disequilibrium of unlinked loci can be used to estimate contemporary effective population size (Ne) of one to a few generations ago and is inflated by about 550 times due to pseudo-replication, highlighting the danger of not handling genetic correlation properly.
Abstract
The Linkage Disequilibrium (LD) of unlinked loci can be used to estimate contemporary effective population size (Ne) of one to a few generations ago. In genomic datasets loci on different chromosomes are considered unlinked, but there are many more pairs of unlinked loci than there are independent pairs of chromosomes, resulting to confidence intervals (C.I.) being too narrow if the non-independence is not taken into account. Simulations were run to investigate the correlation structure among LD of unlinked loci, which can be expressed by the LD of loci along the same chromosomes, based on a discovery of a novel Random Probe LD estimator. We classify the correlation into two categories: overlapping of loci and disjoint pairs. The former is induced from the same locus being considered twice and is the stronger form of correlation. These correlations feed into ρ, a parameter to quantify the degree of pseudo-replication in a dataset, and further a correction formula from which C.I. can be properly inferred. We demonstrate the use of our method via an analysis of genomic data from the malaria-transmitting Anopheles gambiae s.s mosquitoes. Apart from the point and C.I. estimates, we find that is inflated by about 550 times due to pseudo-replication, highlighting the danger of not handling genetic correlation properly.
The linkage disequilibrium (LD) or correlation of alleles at different loci is a fundamental statistic in population genetics. This work presents a new estimator for the standardised LD measure r2 between a pair of loci. The Random Probe (RP) estimator involves generating a set of random dummy loci, before calculating...
Amongst measures of linkage disequilibrium, the correlation measure r2, due originally to Sewall Wright, has several advantages. Although basically a haploid statistic, it can readily be estimated using covariance and correlation from diploid data. Its population expectation can also be found from probability methods....
Pleiotropy, the phenomenon in which a single genetic variant influences multiple traits, is a fundamental feature of genome biology and an important consideration in statistical genetics. Many statistical methods exist that leverage pleiotropy to increase the power of association tests, but they are often designed to i...
Eva Biswas, N. Chatterjee, Zhe-Yu Wang· 0 citations
Abstract This note evaluates variance- and heterozygosity-based genomic estimates of F-statistics through algebraic derivation and simulation. Genetic variance was partitioned among populations, among individuals within populations, and within individuals. The method-of-moments estimator of FIS was evaluated in 100 Mon...
E. Tambarussi· Crop Breeding and Applied Bi...· 0 citations
Genomic epidemiology has several methods for identifying which sampled individuals are linked by close proximity in the transmission chain, whether by grouping them into clusters under a genetic distance threshold or by identifying probable direct transmission pairs. Individuals found to have no sampled neighbours---si...
M. Hall· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.