Spark-Based Label Diffusion and Label Selection Community Detection Algorithm for Metagenome Sequence Clustering
-
Published:2023-11-07
Issue:1
Volume:16
Page:
-
ISSN:1875-6883
-
Container-title:International Journal of Computational Intelligence Systems
-
language:en
-
Short-container-title:Int J Comput Intell Syst
Author:
Wu Zhengjiang,Wu Xuyang,Luo Junwei
Abstract
AbstractIt is a challenge to assemble an enormous amount of metagenome data in metagenomics. Usually, metagenome cluster sequence before assembly accelerates the whole process. In SpaRC, sequences are defined as nodes and clustered by a parallel label propagation algorithm (LPA). To address the randomness of label selection from the parallel LPA during clustering and improve the completeness of metagenome sequence clustering, Spark-based parallel label diffusion and label selection community detection algorithm is proposed in the paper to obtain more accurate clustering results. In this paper, the importance of sequence is defined based on the Jaccard similarity coefficient and its degree. The core sequence is defined as the one with the largest importance in its located community. Three strategies are formulated to reduce the randomness of label selection. Firstly, the core sequence label diffuses over its located cluster and becomes the initial label of other sequences. Those sequences that do not receive an initial label will select the sequence label with the highest importance in the neighbor sequences. Secondly, we perform improved label propagation in order of label frequency and sequence importance to reduce the randomness of label selection. Finally, a merge small communities step is added to increase the completeness of clustered clusters. The experimental results show that our proposed algorithm can effectively reduce the randomness of label selection, improve the purity, completeness, and F-Measure and reduce the runtime of metagenome sequence clustering.
Funder
National Natural Science Foundation of China Innovative and Scientific Research Team of Henan Polytechnic University
Publisher
Springer Science and Business Media LLC
Subject
Computational Mathematics,General Computer Science
Reference24 articles.
1. Yunyan, Z., Min, L., Jiawen, Y.: Recovering metagenome-assembled genomes from shotgun metagenomic sequencing data: methods, applications, challenges, and opportunities. Microbiol. Res. 260, 127 (2022) 2. Wentao, Z., Fuhan, Y., Shiyu, M., Ruiliang, W., Haotian, C., Yuefei, R., Shenghua, L., Pengfei, W., Yang, Y., Wei, L., Junfeng, Z., Xudong, Y.: Bladder cancer-associated microbiota: recent advances and future perspectives. Heliyon 9(1), e13012 (2023) 3. Fadiji, A.E., Babalola, O.O.: Metagenomics methods for the study of plant-associated microbial communities: a review. J. Microbiol. Methods 170(2), 105 (2020) 4. Wang, F.Y., Qin, R., Wang, X., Hu, B.: Metasocieties in metaverse: metaeconomics and metamanagement for metaenterprises and metacities. IEEE Trans. Comput. Soc. Syst. 9(1), 2–7 (2022) 5. Kévin, V., Pierre, M., Maud, T., Jean-Baptiste, V., Jean-Philippe, V.: Large-scale machine learning for metagenomics sequence classification. Bioinformatics (Oxford, England) (2016). https://doi.org/10.1093/bioinformatics/btv683
|
|