Abstract
AbstractHuman genetic diversity remains to be better understood. We here analyzed data from the 1000 Genomes Project and defined group specific fixed alleles (GSFAs) as those that are likely fixed in one ethnic group but non-fixed in at least one other group. The fraction of derived alleles in GSFAs indicates relative distance to apes because such alleles are absent in apes. Our results show that different groups differed in GSFA numbers consistent with known genetic diversity patterns, but also differed in the fraction of derived alleles in GSFAs throughout the entire genome, with East Asians having the largest fraction, followed by South Asians, Europeans, Native Americans, and Africans. Fast evolving sites such as intergenic regions were enriched with derived alleles and showed greater differences in GSFA numbers between East Asians and Africans. Furthermore, GSFAs in East Asians are mostly not fixed in other groups especially Africans, which was particularly more pronounced for fast evolving noncoding variants, while GSFAs in Africans are mostly also fixed in East Asians. Finally, variants that are likely non-neutral such as those leading to stop codon gain/loss and splice donor/acceptor gain/loss showed patterns similar to those of fast-evolving noncoding variants. These results can be accounted for by the maximum genetic diversity theory but not by the neutral theory or its inference that Eurasians suffered bottlenecks, and have implications for better management of group specific genetic diseases.
Publisher
Cold Spring Harbor Laboratory