Affiliation:
1. Hunan Engineering & Technology Research Center for Agricultural Big Data Analysis & Decision-Making, College of Plant Protection, Hunan Agricultural University, Changsha 410128, China
Abstract
Germplasm identification is essential for plant breeding and conservation. In this study, we developed a new method, DT-PICS, for efficient and cost-effective SNP selection in germplasm identification. The method, based on the decision tree concept, could efficiently select the most informative SNPs for germplasm identification by recursively partitioning the dataset based on their overall high PIC values, instead of considering individual SNP features. This method reduces redundancy in SNP selection and enhances the efficiency and automation of the selection process. DT-PICS demonstrated significant advantages in both the training and testing datasets and exhibited good performance on independent prediction, which validates its effectiveness. Thirteen simplified SNP sets were extracted from 749,636 SNPs in 1135 Arabidopsis varieties resequencing datasets, including a total of 769 DT-PICS SNPs, with an average of 59 SNPs per set. Each simplified SNP set could distinguish between the 1135 Arabidopsis varieties. Simulations demonstrated that using a combination of two simplified SNP sets for identification can effectively increase the fault tolerance in independent validation. In the testing dataset, two potentially mislabeled varieties (ICE169 and Star-8) were identified. For 68 same-named varieties, the identification process achieved 94.97% accuracy and only 30 shared markers on average; for 12 different-named varieties, the germplasm to be tested could be effectively distinguished from 1,134 other varieties while grouping extremely similar varieties (Col-0) together, reflecting their actual genetic relatedness. The results suggest that the DT-PICS provides an efficient and accurate approach to SNP selection in germplasm identification and management, offering strong support for future plant breeding and conservation efforts.
Funder
Special Funds for Construction of Innovative Provinces in Hunan Province
Key Research and Development Program of Hubei Province
Natural Science Foundation of Hunan Province
Open Research Fund of State Key Laboratory of Hybrid Rice
Wuhan University
Hunan University Student Innovation and Entrepreneurship Training Program
Subject
Inorganic Chemistry,Organic Chemistry,Physical and Theoretical Chemistry,Computer Science Applications,Spectroscopy,Molecular Biology,General Medicine,Catalysis
Reference28 articles.
1. Current status of the multinational Arabidopsis community;Parry;Plant Direct,2020
2. Verification of Arabidopsis stock collectionpsis stock collections using SNPmatch, a tool for genotyping high-plexed samples;Pisupati;Sci. Data,2017
3. DNA fingerprinting and new tools for fine-scale discrimination of Arabidopsis thaliana accessions;Simon;Plant J. Cell Mol. Biol.,2012
4. El Bakkali, A., Essalouh, L., Tollon, C., Rivallan, R., Mournet, P., Moukhli, A., Zaher, H., Mekkaoui, A., Hadidou, A., and Sikaoui, L. (2019). Characterization of worldwide olive germplasm banks of Marrakech (Morocco) and Córdoba (Spain): Towards management and use of olive germplasm in breeding programs. PLoS ONE, 14.
5. Molecular markers for characterization and conservation of plant genetic resources;Dar;Indian J. Agric. Sci.,2019
Cited by
2 articles.
订阅此论文施引文献
订阅此论文施引文献,注册后可以免费订阅5篇论文的施引文献,订阅后可以查看论文全部施引文献