Classification and selection of the main features for the identification of toxicity in Agaricus and Lepiota with machine learning algorithms

Author:

Ortiz-Letechipia Jacqueline S.1,Galvan-Tejada Carlos E.1,Galván-Tejada Jorge I.1,Soto-Murillo Manuel A.1,Acosta-Cruz Erika2,Gamboa-Rosales Hamurabi1,Celaya Padilla José María1,Luna-García Huizilopoztli1

Affiliation:

1. Unidad Académica de Ingeniería Eléctrica, Universidad Autónoma de Zacatecas, Zacatecas, Zacatecas, México

2. Departamento de Biotecnología., Universidad Autónoma de Coahuila, Saltillo, Coahuila, México

Abstract

The occurrence of fungi is cosmopolitan, and while some mushroom species are beneficial to human health, others can be toxic and cause illness problems. This study aimed to analyze the organoleptic, ecological, and morphological characteristics of a group of fungal specimens and identify the most significant features to develop models for fungal toxicity classification using genetic algorithms and LASSO regression. The results of the study indicated that odor, spore print color, and habitat were the most significant characteristics identified by the genetic algorithm GALGO. Meanwhile, odor, gill size, stalk shape, and twelve other features were the relevant characteristics identified by LASSO regression. The importance score of the odor variable was 99.99%, gill size obtained 73.7%, stalk shape scored 39.9%, and the remaining variables did not score higher than 18%. Logistic regression, k-nearest neighbor (KNN), and XG-Boost classification algorithms were used to develop models using the features selected by both GALGO and LASSO. The models were evaluated using sensitivity, specificity, and accuracy metrics. The models with the highest AUC values were XGBoost, with a maximum value of 0.99 using the features selected by LASSO, followed by KNN with a maximum value of 0.99. The GALGO selection resulted in a maximum AUC of 0.98 in KNN and XGBoost. The models developed in this study have the potential to aid in the accurate identification of toxic fungi, which can prevent health problems caused by their consumption.

Publisher

PeerJ

Reference34 articles.

1. Prediction of whether mushroom is edible or poisonous using back-propagation neural network;Alkronz;International Journal of Corpus Linguistics,2019

2. Minería de Datos;Ballesteros;RECIMUNDO,2018

3. Mining educational data to analyze students’ performance;Baradwaj;International Journal of Advanced Computer Science and Applications,2012

4. Incorporating domain knowledge in machine learning for soccer outcome prediction;Berrar;Machine Learning,2019

5. k-Nearest neighbour classifiers—a tutorial;Cunningham;ACM Computing Surveys,2021

同舟云学术

1.学者识别学者识别

2.学术分析学术分析

3.人才评估人才评估

"同舟云学术"是以全球学者为主线,采集、加工和组织学术论文而形成的新型学术文献查询和分析系统,可以对全球学者进行文献检索和人才价值评估。用户可以通过关注某些学科领域的顶尖人物而持续追踪该领域的学科进展和研究前沿。经过近期的数据扩容,当前同舟云学术共收录了国内外主流学术期刊6万余种,收集的期刊论文及会议论文总量共计约1.5亿篇,并以每天添加12000余篇中外论文的速度递增。我们也可以为用户提供个性化、定制化的学者数据。欢迎来电咨询!咨询电话:010-8811{复制后删除}0370

www.globalauthorid.com

TOP

Copyright © 2019-2024 北京同舟云网络信息技术有限公司
京公网安备11010802033243号  京ICP备18003416号-3