Pangenome Graph Construction from Genome Alignment with Minigraph-Cactus-Reference-Cited by-同舟云学术

Pangenome Graph Construction from Genome Alignment with Minigraph-Cactus

Published:2022-10-07 Issue: Volume: Page:
ISSN:
Container-title:
language:
Short-container-title:

Author:

Hickey Glenn^ORCID,Monlong Jean^ORCID,Ebler Jana^ORCID,Novak Adam^ORCID,Eizenga Jordan M.^ORCID,Gao Yan,Marschall Tobias^ORCID,Li Heng^ORCID,Paten Benedict^ORCID,

Abstract

AbstractReference genomes provide mapping targets and coordinate systems but introduce biases when samples under study diverge sufficiently from them. Pangenome references seek to address this by storing a representative set of diverse haplotypes and their alignment, usually as a graph. Alternate alleles determined by variant callers can be used to construct pangenome graphs, but thanks to advances in long-read sequencing, high-quality phased assemblies are becoming widely available. Constructing a pangenome graph directly from assemblies, as opposed to variant calls, leverages the graph’s ability to consistently represent variation at different scales and reduces biases introduced by reference-based variant calls. Pangenome construction in this way is equivalent to multiple genome alignment. Here we present the Minigraph-Cactus pangenome pipeline, a method to create pangenomes directly from whole-genome alignments, and demonstrate its ability to scale to 90 human haplotypes from the Human Pangenome Reference Consortium (HPRC). This tool was designed to build graphs containing all forms of genetic variation while still being practical for use with current mapping and genotyping tools. We show that this graph is useful both for studying variation within the input haplotypes, but also as a basis for achieving state of the art performance in short and long read mapping, small variant calling and structural variant genotyping. We further measure the effect of the quality and completeness of reference genomes used for analysis within the pangenomes, and show that using the CHM13 reference from the Telomere-to-Telomere Consortium improves the accuracy of our methods, even after projecting back to GRCh38. We also demonstrate that our method can apply to nonhuman data by showing improved mapping and variant detection sensitivity with aDrosophila melanogasterpangenome.

Publisher

Cold Spring Harbor Laboratory

Reference56 articles.

1. Pangenome Graphs

2. The Need for a Human Pangenome Reference Sequence;Annu Rev Genomics Hum Genet,2021

3. Variation graph toolkit improves read mapping by representing genetic variation in the reference

4. Mapping and characterization of structural variation in 17,795 human genomes

5. Hickey G , Heller D , Monlong J , Sibbesen JA , Sirén J , Eizenga J et al. Genotyping structural variants in pangenome graphs using the vg toolkit. Genome Biol 2020; 21: 35.

Cited by 25 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Allopolyploidy expanded gene content but not pangenomic variation in the hexaploid oilseedCamelina sativa;2024-08-16

2. A Draft Pacific Ancestry Pangenome Reference;2024-08-09

3. PPanG: a precision pangenome browser enabling nucleotide-level analysis of genomic variations in individual genomes and their graph-based pangenome;BMC Genomics;2024-04-24

4. Pan-chloroplast genomes for accession-specific marker development in Hibiscus syriacus;Scientific Data;2024-02-27

5. Antibiotic-Free Gene Vectors: A 25-Year Journey to Clinical Trials;Genes;2024-02-20