Asking the right questions for mutagenicity prediction from BioMedical text-Reference-Cited by-同舟云学术

Asking the right questions for mutagenicity prediction from BioMedical text

Published:2023-12-18 Issue:1 Volume:9 Page:
ISSN:2056-7189
Container-title:npj Systems Biology and Applications
language:en
Short-container-title:npj Syst Biol Appl

Author:

Acharya Sathwik^ORCID,Shinada Nicolas K.,Koyama Naoki,Ikemori Megumi,Nishioka Tomoki,Hitaoka Seiji,Hakura Atsushi,Asakura Shoji,Matsuoka Yukiko,Palaniappan Sucheendra K.^ORCID

Abstract

AbstractAssessing the mutagenicity of chemicals is an essential task in the drug development process. Usually, databases and other structured sources for AMES mutagenicity exist, which have been carefully and laboriously curated from scientific publications. As knowledge accumulates over time, updating these databases is always an overhead and impractical. In this paper, we first propose the problem of predicting the mutagenicity of chemicals from textual information in scientific publications. More simply, given a chemical and evidence in the natural language form from publications where the mutagenicity of the chemical is described, the goal of the model/algorithm is to predict if it is potentially mutagenic or not. For this, we first construct a golden standard data set and then propose MutaPredBERT, a prediction model fine-tuned on BioLinkBERT based on a question-answering formulation of the problem. We leverage transfer learning and use the help of large transformer-based models to achieve a Macro F1 score of >0.88 even with relatively small data for fine-tuning. Our work establishes the utility of large language models for the construction of structured sources of knowledge bases directly from scientific publications.

Funder

United States Department of Defense | United States Navy | ONR | Office of Naval Research Global

Publisher

Springer Science and Business Media LLC

Subject

Applied Mathematics,Computer Science Applications,Drug Discovery,General Biochemistry, Genetics and Molecular Biology,Modeling and Simulation

Link

https://www.nature.com/articles/s41540-023-00324-2.pdf

Reference37 articles.

1. Stead, A. G., Hasselblad, V., Creason, J. P. & Claxton, L. Modeling the Ames test. Mutation Res./Environ. Mutagenesis Relat. Subjects 85, 13–27 (1981).

2. Nantasenamat, C., Isarankura-Na-Ayudhya, C., Naenna, T. & Prachayasittikul, V. A Practical Overview of Quantitative Structure-Activity Relationship. EXCLI J. 8, 74–88. https://doi.org/10.17877/DE290R-690 (2009).

3. Shinada, N. K. et al. Optimizing machine-learning models for mutagenicity prediction through better feature selection. Mutagenesis (2022).

4. Mayr, A., Klambauer, G., Unterthiner, T. & Hochreiter, S. Deeptox: toxicity prediction using deep learning. Front. Environ. Sci. 3, 80 (2016).

5. Lin, Z. et al. Evolutionary-scale prediction of atomic level protein structure with a language model. bioRxiv. https://www.biorxiv.org/content/early/2022/10/31/2022.07.20.500902 (2022).

Cited by 1 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Comparative in silico analysis of CNS-active molecules targeting the blood–brain barrier choline transporter for Alzheimer’s disease therapy;In Silico Pharmacology;2024-07-31