AllTheBacteria - all bacterial genomes assembled, available and searchable-Reference-Cited by-同舟云学术

AllTheBacteria - all bacterial genomes assembled, available and searchable

Published:2024-03-11 Issue: Volume: Page:
ISSN:
Container-title:
language:
Short-container-title:

Author:

Hunt Martin^ORCID,Lima Leandro^ORCID,Shen Wei^ORCID,Lees John^ORCID,Iqbal Zamin^ORCID

Abstract

AbstractThe bacterial sequence data publicly available at the global DNA archives is a vast source of information on the evolution of bacteria and their mobile elements. However, most of it is either unassembled or inconsistently assembled and QC-ed. This makes it unsuitable for large-scale analyses, and inaccessible for most researchers to use. In 2021 Blackwell et al therefore released a uniformly assembled set of 661,405 genomes, consisting of all publicly available whole genome sequenced bacterial isolate data as of November 2018, along with various search indexes. In this study we extend that dataset by 4.5 years (up to May 2023), tripling the number of genomes. We also expand the scope, as we begin a global collaborative project to generate annotations for different species as desired by different research communities.In this study we describe the initial v0.1 data release of 1,932,812 assemblies (combining 1,271,428 new assemblies with the 661k dataset). All 1.9 million have been uniformly re-processed for quality criteria and to give taxonomic abundance estimates with respect to the GTDB phylogeny. Using an evolution-informed compression approach, the full set of genomes is just 102Gb in batched xz archives. We also provide multiple search indexes. Finally, we outline plans for future annotations to be provided in further releases.

Publisher

Cold Spring Harbor Laboratory

Reference21 articles.

1. Exploring bacterial diversity via a curated and searchable snapshot of archived DNA sequences

2. Large-scale sequence comparisons with sourmash;F1000Research,2019

3. Fast and flexible bacterial genomic epidemiology with PopPUNK

4. Timo Bingmann , Phelim Bradley , Florian Gauger , and Zamin Iqbal . COBS: a Compact Bit-Sliced Signature Index. 2019.

5. Genomic epidemiology reveals multidrug resistant plasmid spread between Vibrio cholerae lineages in Yemen;Nature Microbiology,2023

Cited by 7 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. LexicMap: efficient sequence alignment against millions of prokaryotic genomes;2024-08-31

2. PanKB: An interactive microbial pangenome knowledgebase for research, biotechnological innovation, and knowledge mining;2024-08-19

3. In vivo selection of carbapenem resistance during persistent Klebsiella pneumoniae sequence type 395 bloodstream infection due to OmpK36 deletion;Antimicrobial Agents and Chemotherapy;2024-08-07

4. Logan: Planetary-Scale Genome Assembly Surveys Life’s Diversity;2024-07-31

5. ParallelEvolCCM: Quantifying co-evolutionary patterns among genomic features;2024-06-14