PubMed Text Similarity Model and its application to curation efforts in the Conserved Domain Database-Reference-Cited by-同舟云学术

PubMed Text Similarity Model and its application to curation efforts in the Conserved Domain Database

Published:2019-01-01 Issue: Volume:2019 Page:
ISSN:1758-0463
Container-title:Database
language:en
Short-container-title:

Author:

Islamaj Rezarta¹,Wilbur W John¹,Xie Natalie¹,Gonzales Noreen R¹,Thanki Narmada¹,Yamashita Roxanne¹,Zheng Chanjuan¹,Marchler-Bauer Aron¹,Lu Zhiyong¹

Affiliation:

1. National Center for Biotechnology Information, National Library of Medicine, National Institutes of Health, Bethesda, MD, 20894, USA

Abstract

AbstractThis study proposes a text similarity model to help biocuration efforts of the Conserved Domain Database (CDD). CDD is a curated resource that catalogs annotated multiple sequence alignment models for ancient domains and full-length proteins. These models allow for fast searching and quick identification of conserved motifs in protein sequences via Reverse PSI-BLAST. In addition, CDD curators prepare summaries detailing the function of these conserved domains and specific protein families, based on published peer-reviewed articles. To facilitate information access for database users, it is desirable to specifically identify the referenced articles that support the assertions of curator-composed sentences. Moreover, CDD curators desire an alert system that scans the newly published literature and proposes related articles of relevance to the existing CDD records. Our approach to address these needs is a text similarity method that automatically maps a curator-written statement to candidate sentences extracted from the list of referenced articles, as well as the articles in the PubMed Central database. To evaluate this proposal, we paired CDD description sentences with the top 10 matching sentences from the literature, which were given to curators for review. Through this exercise, we discovered that we were able to map the articles in the reference list to the CDD description statements with an accuracy of 77%. In the dataset that was reviewed by curators, we were able to successfully provide references for 86% of the curator statements. In addition, we suggested new articles for curator review, which were accepted by curators to be added into the reference list at an acceptance rate of 50%. Through this process, we developed a substantial corpus of similar sentences from biomedical articles on protein sequence, structure and function research, which constitute the CDD text similarity corpus. This corpus contains 5159 sentence pairs judged for their similarity on a scale from 1 (low) to 5 (high) doubly annotated by four CDD curators. Curator-assigned similarity scores have a Pearson correlation coefficient of 0.70 and an inter-annotator agreement of 85%. To date, this is the largest biomedical text similarity resource that has been manually judged, evaluated and made publicly available to the community to foster research and development of text similarity algorithms.

Publisher

Oxford University Press (OUP)

Subject

General Agricultural and Biological Sciences,General Biochemistry, Genetics and Molecular Biology,Information Systems

Link

http://academic.oup.com/database/article-pdf/doi/10.1093/database/baz064/28895934/baz064.pdf

Reference34 articles.

1. Manual curation is not sufficient for annotation of genomic databases;Baumgartner;Bioinformatics,2007

2. Perspective: sustaining the big-data ecosystem;Bourne;Nature,2015

3. On expert curation and scalability: UniProtKB/Swiss-Prot as a case study;Poux;Bioinformatics,2017

4. Text mining for the biocuration workflow;Hirschman;Database (Oxford),2012

5. Evaluation of text-mining systems for biology: overview of the second BioCreative community challenge;Krallinger;Genome Biol.,2008

Cited by 10 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Enhancing query relevance: leveraging SBERT and cosine similarity for optimal information retrieval;International Journal of Speech Technology;2024-08-16

2. Information extraction and application for constructing guidance corpus of welding fabrication;Proceedings of the Institution of Mechanical Engineers, Part B: Journal of Engineering Manufacture;2023-01-11

3. A comparative evaluation of biomedical similar article recommendation;Journal of Biomedical Informatics;2022-07

4. BCDRRLE: A Bidirectional Cross-Dynamic Round Robin Learning Encoder Model for Medical Sentence Similarity (Preprint);2022-03-20

5. Genome-wide analysis of potassium transport genes in Gossypium raimondii suggest a role of GrHAK/KUP/KT8, GrAKT2.1 and GrAKT1.1 in response to abiotic stress;Plant Physiology and Biochemistry;2022-01