A decision-theoretic approach to the evaluation of machine learning algorithms in computational drug discovery-Reference-Cited by-同舟云学术

A decision-theoretic approach to the evaluation of machine learning algorithms in computational drug discovery

Published:2019-05-09 Issue:22 Volume:35 Page:4656-4663
ISSN:1367-4803
Container-title:Bioinformatics
language:en
Short-container-title:

Author:

Watson Oliver P¹,Cortes-Ciriano Isidro¹²,Taylor Aimee R³⁴,Watson James A⁵⁶^ORCID

Affiliation:

1. Goring on Thames, Evariste Technologies Ltd., RG8 9AL UK

2. Department of Chemistry, Centre for Molecular Science Informatics, University of Cambridge, Lensfield Road, Cambridge CB2 1EW, UK

3. Department of Epidemiology, Center for Communicable Disease Dynamics, Harvard T.H. Chan School of Public Health, Boston, MA 02115 USA

4. Infectious Disease Microbiome Program, Broad Institute, Cambridge, MA 02142 USA

5. Nuffield Department of Medicine, Centre for Tropical Medicine and Global Health, University of Oxford, Oxford OX3, 7LF UK

6. Mahidol-Oxford Tropical Medicine Research Unit, Faculty of Tropical Medicine, Mahidol University, Bangkok 10400, Thailand

Abstract

Abstract Motivation Artificial intelligence, trained via machine learning (e.g. neural nets, random forests) or computational statistical algorithms (e.g. support vector machines, ridge regression), holds much promise for the improvement of small-molecule drug discovery. However, small-molecule structure-activity data are high dimensional with low signal-to-noise ratios and proper validation of predictive methods is difficult. It is poorly understood which, if any, of the currently available machine learning algorithms will best predict new candidate drugs. Results The quantile-activity bootstrap is proposed as a new model validation framework using quantile splits on the activity distribution function to construct training and testing sets. In addition, we propose two novel rank-based loss functions which penalize only the out-of-sample predicted ranks of high-activity molecules. The combination of these methods was used to assess the performance of neural nets, random forests, support vector machines (regression) and ridge regression applied to 25 diverse high-quality structure-activity datasets publicly available on ChEMBL. Model validation based on random partitioning of available data favours models that overfit and ‘memorize’ the training set, namely random forests and deep neural nets. Partitioning based on quantiles of the activity distribution correctly penalizes extrapolation of models onto structurally different molecules outside of the training data. Simpler, traditional statistical methods such as ridge regression can outperform state-of-the-art machine learning methods in this setting. In addition, our new rank-based loss functions give considerably different results from mean squared error highlighting the necessity to define model optimality with respect to the decision task at hand. Availability and implementation All software and data are available as Jupyter notebooks found at https://github.com/owatson/QuantileBootstrap. Supplementary information Supplementary data are available at Bioinformatics online.

Publisher

Oxford University Press (OUP)

Subject

Computational Mathematics,Computational Theory and Mathematics,Computer Science Applications,Molecular Biology,Biochemistry,Statistics and Probability

Link

http://academic.oup.com/bioinformatics/advance-article-pdf/doi/10.1093/bioinformatics/btz293/28670671/btz293.pdf

Reference41 articles.

1. Can we learn to distinguish between ‘drug-like’ and ‘nondrug-like’ molecules?;Ajay;J. Med. Chem,1998

2. Optimal classifier selection and negative bias in error rate estimation: an empirical study on high-dimensional prediction;Boulesteix;BMC Med. Res. Methodol,2009

3. Cross-validation under separate sampling: strong bias and how to correct it;Braga-Neto;Bioinformatics,2014

4. Drug design by machine learning: support vector machines for pharmaceutical data analysis;Burbidge;Comput. Chem,2001

Cited by 13 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Extrapolation is not the same as interpolation;Machine Learning;2024-07-23

2. AI's role in pharmaceuticals: Assisting drug design from protein interactions to drug development;Artificial Intelligence Chemistry;2024-06

3. Future Directions and Challenges in Overcoming Drug Resistance in Cancer;Drug Resistance in Cancer: Mechanisms and Strategies;2024

4. Graph Neural Tree: A novel and interpretable deep learning-based framework for accurate molecular property predictions;Analytica Chimica Acta;2023-03

5. Machine learning in metastatic cancer research: Potentials, possibilities, and prospects;Computational and Structural Biotechnology Journal;2023