Toward a Coronavirus Knowledge Graph-Reference-Cited by-同舟云学术

Toward a Coronavirus Knowledge Graph

Published:2021-06-29 Issue:7 Volume:12 Page:998
ISSN:2073-4425
Container-title:Genes
language:en
Short-container-title:Genes

Author:

Zhang Peng,Bu Yi^ORCID,Jiang Peng,Shi Xiaowen,Lun Bing,Chen Chongyan^ORCID,Syafiandini Arida Ferti,Ding Ying,Song Min

Abstract

This study builds a coronavirus knowledge graph (KG) by merging two information sources. The first source is Analytical Graph (AG), which integrates more than 20 different public datasets related to drug discovery. The second source is CORD-19, a collection of published scientific articles related to COVID-19. We combined both chemo genomic entities in AG with entities extracted from CORD-19 to expand knowledge in the COVID-19 domain. Before populating KG with those entities, we perform entity disambiguation on CORD-19 collections using Wikidata. Our newly built KG contains at least 21,700 genes, 2500 diseases, 94,000 phenotypes, and other biological entities (e.g., compound, species, and cell lines). We define 27 relationship types and use them to label each edge in our KG. This research presents two cases to evaluate the KG’s usability: analyzing a subgraph (ego-centered network) from the angiotensin-converting enzyme (ACE) and revealing paths between biological entities (hydroxychloroquine and IL-6 receptor; chloroquine and STAT1). The ego-centered network captured information related to COVID-19. We also found significant COVID-19-related information in top-ranked paths with a depth of three based on our path evaluation.

Funder

National Research Foundation of Korea

National Science Foundation in the United States

Publisher

MDPI AG

Subject

Genetics (clinical),Genetics

Link

https://www.mdpi.com/2073-4425/12/7/998/pdf

Reference55 articles.

1. Coronavirus Disease (COVID-19)https://www.who.int/emergencies/diseases/novel-coronavirus-2019

2. A Bibliometric Analysis of COVID-19 Research Activity: A Call for Increased Output