MetaConClust - Unsupervised Binning of Metagenomics Data using Consensus Clustering-Reference-Cited by-同舟云学术

MetaConClust - Unsupervised Binning of Metagenomics Data using Consensus Clustering

Published:2022-02 Issue:2 Volume:23 Page:137-146
ISSN:1389-2029
Container-title:Current Genomics
language:en
Short-container-title:CG

Author:

Sharma Anu¹,Sinha Dipro²,Mishra Dwijesh Chandra¹,Rai Anil¹,Lal Shashi Bhushan¹,Kumar Sanjeev¹,Farooqi Moh. Samir¹,Chaturvedi Krishna Kumar¹

Affiliation:

1. Division of Agriculture Bioinformatics, ICARIASRI, New Delhi- 110012, India

2. Research Scholar, PG School, ICAR-IARI, New Delhi-110012, India

Abstract

Background: Binning of metagenomic reads is an active area of research, and many unsupervised machine learning-based techniques have been used for taxonomic independent binning of metagenomic reads. Objective: It is important to find the optimum number of the cluster as well as develop an efficient pipeline for deciphering the complexity of the microbial genome. Method: Applying unsupervised clustering techniques for binning requires finding the optimal number of clusters beforehand and is observed to be a difficult task. This paper describes a novel method, MetaConClust, using coverage information for grouping of contigs and automatically finding the optimal number of clusters for binning of metagenomics data using a consensus-based clustering approach. The coverage of contigs in a metagenomics sample has been observed to be directly proportional to the abundance of species in the sample and is used for grouping of data in the first phase by MetaConClust. The Partitioning Around Medoid (PAM) method is used for clustering in the second phase for generating bins with the initial number of clusters determined automatically through a consensus-based method. Results: Finally, the quality of the obtained bins is tested using silhouette index, rand Index, recall, precision, and accuracy. Performance of MetaConClust is compared with recent methods and tools using benchmarked low complexity simulated and real metagenomic datasets and is found better for unsupervised and comparable for hybrid methods. Conclusion: This is suggestive of the proposition that the consensus-based clustering approach is a promising method for automatically finding the number of bins for metagenomics data.

Publisher

Bentham Science Publishers Ltd.

Subject

Genetics (clinical),Genetics

Reference28 articles.

1. Handelsman J.; Metagenomics: Application of genomics to uncultured microorganisms. Microbiol Mol Biol Rev 2004,68(4),669-685

2. Meyer F.; Paarmann D.; D’Souza M.; Olson R.; Glass E.M.; Kubal M.; Paczian T.; Rodriguez A.; Stevens R.; Wilke A.; Wilkening J.; Edwards R.A.; The metagenomics RAST server - a public resource for the automatic phylogenetic and functional analysis of metagenomes. BMC Bioinformatics 2008,9(1),386-393

3. Sedlar K.; Kupkova K.; Provaznik I.; Bioinformatics strategies for taxonomy independent binning and visualization of sequences in shotgun metagenomics. Comput Struct Biotechnol J 2016,15,48-55

4. Huson D.H.; Auch A.F.; Qi J.; Schuster S.C.; MEGAN analysis of metagenomic data. Genome Res 2007,17(3),377-386

5. Segata N.; Waldron L.; Ballarini A.; Narasimhan V.; Jousson O.; Huttenhower C.; Metagenomic microbial community profiling using unique clade-specific marker genes. Nat Methods 2012,9(8),811-814

Cited by 4 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Metagenomic approaches and opportunities in arid soil research;Science of The Total Environment;2024-11

2. MethSemble-6mA: an ensemble-based 6mA prediction server and its application on promoter region of LBD gene family in Poaceae;Frontiers in Plant Science;2023-10-09

3. EpiSemble: A Novel Ensemble-based Machine-learning Framework for Prediction of DNA N6-methyladenine Sites Using Hybrid Features Selection Approach for Crops;Current Bioinformatics;2023-08

4. A Deep Clustering-based Novel Approach for Binning of Metagenomics Data;Current Genomics;2022-08