Towards end-to-end disease prediction from raw metagenomic data-Reference-Cited by-同舟云学术

Towards end-to-end disease prediction from raw metagenomic data

Published:2020-10-29 Issue: Volume: Page:
ISSN:
Container-title:
language:
Short-container-title:

Author:

Queyrel Maxence,Prifti Edi,Templier Alexandre,Zucker Jean-Daniel

Abstract

ABSTRACTAnalysis of the human microbiome using metagenomic sequencing data has demonstrated high ability in discriminating various human diseases. Raw metagenomic sequencing data require multiple complex and computationally heavy bioinformatics steps prior to data analysis. Such data contain millions of short sequences read from the fragmented DNA sequences and are stored as fastq files. Conventional processing pipelines consist in multiple steps including quality control, filtering, alignment of sequences against genomic catalogs (genes, species, taxonomic levels, functional pathways, etc.). These pipelines are complex to use, time consuming and rely on a large number of parameters that often provide variability and impact the estimation of the microbiome elements.Training Deep Neural Networks directly from raw sequencing data is a promising approach to bypass some of the challenges associated with mainstream bioinformatics pipelines. Most of these methods use the concept of word and sentence embeddings that create a meaningful and numerical representation of DNA sequences, while extracting features and reducing the dimensionality of the data. In this paper we present an end-to-end approach that classifies patients into disease groups directly from raw metagenomic reads: metagenome2vec. This approach is composed of four steps (i) generating a vocabulary of k-mers and learning their numerical embeddings; (ii) learning DNA sequence (read) embeddings; (iii) identifying the genome from which the sequence is most likely to come and (iv) training a multiple instance learning classifier which predicts the phenotype based on the vector representation of the raw data. An attention mechanism is applied in the network so that the model can be interpreted, assigning a weight to the influence of the prediction for each genome. Using two public real-life data-sets as well a simulated one, we demonstrated that this original approach reaches high performance, comparable with the state-of-the-art methods applied directly on processed data though mainstream bioinformatics workflows. These results are encouraging for this proof of concept work. We believe that with further dedication, the DNN models have the potential to surpass mainstream bioinformatics workflows in disease classification tasks.

Publisher

Cold Spring Harbor Laboratory

Reference51 articles.

1. Arora, Sanjeev , Yingyu Liang , and Tengyu Ma . “A simple but tough-to-beat baseline for sentence embeddings”. In: (2017), p. 16.

2. A Survey of Word Embeddings Evaluation Methods;arXiv:1801.09536 [cs],2018

3. Bengio, Samy and Georg Heigold . “Word Embeddings for Speech Recognition”. In: (2014), p. 5.

4. Enriching Word Vectors with Subword Information;arXiv:1607.04606 [cs],2016

5. Supervised Learning of Universal Sentence Representations from Natural Language Inference Data;arXiv:1705.02364 [cs],2017

Cited by 6 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Deep learning methods in metagenomics: a review;Microbial Genomics;2024-04-17

2. Enhancing Clinical Utility: Utilization of International Standards and Guidelines for Metagenomic Sequencing in Infectious Disease Diagnosis;International Journal of Molecular Sciences;2024-03-15

3. A self-supervised deep learning method for data-efficient training in genomics;Communications Biology;2023-09-11

4. Deep learning methods in metagenomics: a review;2023-08-08

5. PWM2Vec: An Efficient Embedding Approach for Viral Host Specification from Coronavirus Spike Sequences;Biology;2022-03-09