On evaluation metrics for medical applications of artificial intelligence-Reference-Cited by-同舟云学术

On evaluation metrics for medical applications of artificial intelligence

Published:2022-04-08 Issue:1 Volume:12 Page:
ISSN:2045-2322
Container-title:Scientific Reports
language:en
Short-container-title:Sci Rep

Author:

Hicks Steven A.,Strümke Inga,Thambawita Vajira,Hammou Malek,Riegler Michael A.,Halvorsen Pål,Parasa Sravanthi

Abstract

AbstractClinicians and software developers need to understand how proposed machine learning (ML) models could improve patient care. No single metric captures all the desirable properties of a model, which is why several metrics are typically reported to summarize a model’s performance. Unfortunately, these measures are not easily understandable by many clinicians. Moreover, comparison of models across studies in an objective manner is challenging, and no tool exists to compare models using the same performance metrics. This paper looks at previous ML studies done in gastroenterology, provides an explanation of what different metrics mean in the context of binary classification in the presented studies, and gives a thorough explanation of how different metrics should be interpreted. We also release an open source web-based tool that may be used to aid in calculating the most relevant metrics presented in this paper so that other researchers and clinicians may easily incorporate them into their research.

Publisher

Springer Science and Business Media LLC

Subject

Multidisciplinary

Link

https://www.nature.com/articles/s41598-022-09954-8.pdf

Reference24 articles.

1. Nagendran, M. et al. Artificial intelligence versus clinicians: Systematic review of design, reporting standards, and claims of deep learning studies. bmj 368, m689. https://doi.org/10.1136/bmj.m689 (2020).

2. Topol, E. J. High-performance medicine: The convergence of human and artificial intelligence. Nat. Med. 25, 44–56. https://doi.org/10.1038/s41591-018-0300-7 (2019).

3. Schmitz, R. et al. Artificial intelligence in GI endoscopy: Stumbling blocks, gold standards and the role of endoscopy societies. Gut. https://doi.org/10.1136/gutjnl-2020-323115 (2021).

4. Hoogenboom, S. A., Bagci, U. & Wallace, M. B. AI in gastroenterology. The current state of play and the potential. How will it affect our practice and when?. Techn. Gastrointest. Endosc. 22, 150634. https://doi.org/10.1016/j.tgie.2019.150634 (2019).

5. Patel, K. et al. A comparative study on polyp classification using convolutional neural networks. PLOS ONE 15, 1–16. https://doi.org/10.1371/journal.pone.0236452 (2020).

Cited by 263 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Investigation of in silico studies for cytochrome P450 isoforms specificity;Computational and Structural Biotechnology Journal;2024-12

2. Predicting adolescent psychopathology from early life factors: A machine learning tutorial;Global Epidemiology;2024-12

3. Structured adaptive boosting trees for detection of multicellular aggregates in fluorescence intravital microscopy;Microvascular Research;2024-11

4. Immune-based Machine learning Prediction of Diagnosis and Illness State in Schizophrenia and Bipolar Disorder;Brain, Behavior, and Immunity;2024-11

5. A transparent machine learning algorithm uncovers HbA1c patterns associated with therapeutic inertia in patients with type 2 diabetes and failure of metformin monotherapy;International Journal of Medical Informatics;2024-10