Measuring Performance Metrics of Machine Learning Algorithms for Detecting and Classifying Transposable Elements-Reference-Cited by-同舟云学术

Measuring Performance Metrics of Machine Learning Algorithms for Detecting and Classifying Transposable Elements

Published:2020-05-27 Issue:6 Volume:8 Page:638
ISSN:2227-9717
Container-title:Processes
language:en
Short-container-title:Processes

Author:

Orozco-Arias Simon^ORCID,Piña Johan S.,Tabares-Soto Reinel^ORCID,Castillo-Ossa Luis F.,Guyot Romain^ORCID,Isaza Gustavo^ORCID

Abstract

Because of the promising results obtained by machine learning (ML) approaches in several fields, every day is more common, the utilization of ML to solve problems in bioinformatics. In genomics, a current issue is to detect and classify transposable elements (TEs) because of the tedious tasks involved in bioinformatics methods. Thus, ML was recently evaluated for TE datasets, demonstrating better results than bioinformatics applications. A crucial step for ML approaches is the selection of metrics that measure the realistic performance of algorithms. Each metric has specific characteristics and measures properties that may be different from the predicted results. Although the most commonly used way to compare measures is by using empirical analysis, a non-result-based methodology has been proposed, called measure invariance properties. These properties are calculated on the basis of whether a given measure changes its value under certain modifications in the confusion matrix, giving comparative parameters independent of the datasets. Measure invariance properties make metrics more or less informative, particularly on unbalanced, monomodal, or multimodal negative class datasets and for real or simulated datasets. Although several studies applied ML to detect and classify TEs, there are no works evaluating performance metrics in TE tasks. Here, we analyzed 26 different metrics utilized in binary, multiclass, and hierarchical classifications, through bibliographic sources, and their invariance properties. Then, we corroborated our findings utilizing freely available TE datasets and commonly used ML algorithms. Based on our analysis, the most suitable metrics for TE tasks must be stable, even using highly unbalanced datasets, multimodal negative class, and training datasets with errors or outliers. Based on these parameters, we conclude that the F1-score and the area under the precision-recall curve are the most informative metrics since they are calculated based on other metrics, providing insight into the development of an ML application.

Publisher

MDPI AG

Subject

Process Chemistry and Technology,Chemical Engineering (miscellaneous),Bioengineering

Link

https://www.mdpi.com/2227-9717/8/6/638/pdf

Reference87 articles.

1. How retrotransposons shape genome regulation

2. Genome-wide analysis of a recently active retrotransposon, Au SINE, in wheat: content, distribution within subgenomes and chromosomes, and gene associations

3. Retrotransposons in Plant Genomes: Structure, Identification, and Classification through Bioinformatics and Machine Learning

4. Structure and Distribution of Centromeric Retrotransposons at Diploid and Allotetraploid Coffea Centromeric and Pericentromeric Regions

5. Assessing genome assembly quality using the LTR Assembly Index (LAI)

Cited by 34 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Optimizing Lung Condition Categorization through a Deep Learning Approach to Chest X-ray Image Analysis;BioMedInformatics;2024-09-10

2. An Essay on Detailed Performance Assessment Ensemble-Based Predictive Modelling in Network Intrusion Detection Systems;2024 IEEE Students Conference on Engineering and Systems (SCES);2024-06-21

3. Integrating Principal Component Analysis and Multi-Input Convolutional Neural Networks for Advanced Skin Lesion Cancer Classification;Applied Sciences;2024-06-17

4. Novel Algorithms for Early Cancer Diagnosis Using Transfer Learning with MobileNetV2 in Thermal Images;KSII Transactions on Internet and Information Systems;2024-03-31

5. Effective Stroke Prediction using Machine Learning Algorithms;Australian Journal of Engineering and Innovative Technology;2024-03-09