MultiStageSearch: a multi-step proteogenomic workflow for taxonomic identification of viral proteome samples adressing database bias

Author:

Pipart Julian,Holstein TanjaORCID,Martens LennartORCID,Muth ThiloORCID

Abstract

AbstractThe recent years, with the global SARS-Cov-2 pandemic, have shown the importance of strain level identification of viral pathogens. While the gold-standard approach for unkown viral sample identification remains genomics, studies have shown the necessity and advantages of orthogonal experimental approaches such as proteomics, based on proteomic database search methods. The databases required as references for both proteins and genome sequences are known to be biased towards certain taxa, such as pathogenic strains or species, or common model organisms. Aditionally, the proteomic databases are not as comprehensive as the genomic databases.We present MultiStageSearch, an iterative database search approach for the taxonomic identification of viral samples combining proteomic and genomic databases. The potentially present species and strains are inferred using a generalist proteomic reference database. MultiStageSearch then automatically creates a proteogenomic database. This database is further pre-processed byfiltering for duplicates as well as clustering of identical ORFs to address potential bias present in the genomic database. Furthermore, the workflow is independent of the strain level NCBI taxonomy, enabling the inference of strains that are not present in the NCBI taxonomy.We performed a benchmark on several viral samples to demonstrate the performance of the strain level taxonomic inference. The benchmark shows superior performance compared to state of the art methods for untargeted strain level inference using proteomic data while being independent of the NCBI taxonomy at strain level.

Publisher

Cold Spring Harbor Laboratory

同舟云学术

1.学者识别学者识别

2.学术分析学术分析

3.人才评估人才评估

"同舟云学术"是以全球学者为主线,采集、加工和组织学术论文而形成的新型学术文献查询和分析系统,可以对全球学者进行文献检索和人才价值评估。用户可以通过关注某些学科领域的顶尖人物而持续追踪该领域的学科进展和研究前沿。经过近期的数据扩容,当前同舟云学术共收录了国内外主流学术期刊6万余种,收集的期刊论文及会议论文总量共计约1.5亿篇,并以每天添加12000余篇中外论文的速度递增。我们也可以为用户提供个性化、定制化的学者数据。欢迎来电咨询!咨询电话:010-8811{复制后删除}0370

www.globalauthorid.com

TOP

Copyright © 2019-2024 北京同舟云网络信息技术有限公司
京公网安备11010802033243号  京ICP备18003416号-3