Discovering drug–target interaction knowledge from biomedical literature-Reference-Cited by-同舟云学术

Discovering drug–target interaction knowledge from biomedical literature

Published:2022-10-07 Issue:22 Volume:38 Page:5100-5107
ISSN:1367-4803
Container-title:Bioinformatics
language:en
Short-container-title:

Author:

Hou Yutai¹^ORCID,Xia Yingce²^ORCID,Wu Lijun²^ORCID,Xie Shufang²,Fan Yang³,Zhu Jinhua³,Qin Tao²,Liu Tie-Yan²

Affiliation:

1. Harbin Institute of Technology , Harbin 150001, China

2. Microsoft Research , Beijing 100080, China

3. University of Science and Technology of China , Hefei 230027, China

Abstract

Abstract Motivation The interaction between drugs and targets (DTI) in human body plays a crucial role in biomedical science and applications. As millions of papers come out every year in the biomedical domain, automatically discovering DTI knowledge from biomedical literature, which are usually triplets about drugs, targets and their interaction, becomes an urgent demand in the industry. Existing methods of discovering biological knowledge are mainly extractive approaches that often require detailed annotations (e.g. all mentions of biological entities, relations between every two entity mentions, etc.). However, it is difficult and costly to obtain sufficient annotations due to the requirement of expert knowledge from biomedical domains. Results To overcome these difficulties, we explore an end-to-end solution for this task by using generative approaches. We regard the DTI triplets as a sequence and use a Transformer-based model to directly generate them without using the detailed annotations of entities and relations. Further, we propose a semi-supervised method, which leverages the aforementioned end-to-end model to filter unlabeled literature and label them. Experimental results show that our method significantly outperforms extractive baselines on DTI discovery. We also create a dataset, KD-DTI, to advance this task and release it to the community. Availability and implementation Our code and data are available at https://github.com/bert-nmt/BERT-DTI. Supplementary information Supplementary data are available at Bioinformatics online.

Publisher

Oxford University Press (OUP)

Subject

Computational Mathematics,Computational Theory and Mathematics,Computer Science Applications,Molecular Biology,Biochemistry,Statistics and Probability

Link

https://academic.oup.com/bioinformatics/advance-article-pdf/doi/10.1093/bioinformatics/btac648/46571897/btac648.pdf

Reference43 articles.

1. Extraction of chemical–protein interactions from the literature using neural networks and narrow instance representation;Antunes;Database,2019

Cited by 6 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. A comprehensive evaluation of large Language models on benchmark biomedical text processing tasks;Computers in Biology and Medicine;2024-03

2. Artificial intelligence generated content (AIGC) in medicine: A narrative review;Mathematical Biosciences and Engineering;2024

3. A Survey of Large Language Models for Healthcare: From Data, Technology, and Applications to Accountability and Ethics;2024

4. A study of generative large language model for medical research and healthcare;npj Digital Medicine;2023-11-16

5. Uncovering the Effects of Genes, Proteins, and Medications on Functions of Wound Healing: A Dependency Rule-Based Text Mining Approach Leveraging GPT-4 based Evaluation;2023 IEEE EMBS International Conference on Biomedical and Health Informatics (BHI);2023-10-15