Abstract
In the biomedical field, there is an ever-increasing number of large, fragmented, and isolated data sources stored in databases and ontologies that use heterogeneous formats and poorly integrated schemes. Researchers and healthcare professionals find it extremely difficult to master this huge amount of data and extract relevant information. In this work, we propose a linked data approach, based on multilayer networks and semantic Web standards, capable of integrating and harmonizing several biomedical datasets with different schemas and semi-structured data through a multi-model database providing polyglot persistence. The domain chosen concerns the analysis and aggregation of available data on neuroendocrine neoplasms (NENs), a relatively rare type of neoplasm. Integrated information includes twelve public datasets available in heterogeneous schemas and formats including RDF, CSV, TSV, SQL, OWL, and OBO. The proposed integrated model consists of six interconnected layers representing, respectively, information on the disease, the related phenotypic alterations, the affected genes, the related biological processes, molecular functions, the involved human tissues, and drugs and compounds that show documented interactions with them. The defined scheme extends an existing three-layer model covering a subset of the mentioned aspects. A client–server application was also developed to browse and search for information on the integrated model. The main challenges of this work concern the complexity of the biomedical domain, the syntactic and semantic heterogeneity of the datasets, and the organization of the integrated model. Unlike related works, multilayer networks have been adopted to organize the model in a manageable and stratified structure, without the need to change the original datasets but by transforming their data “on the fly” to respond to user requests.
Subject
Fluid Flow and Transfer Processes,Computer Science Applications,Process Chemistry and Technology,General Engineering,Instrumentation,General Materials Science
Reference38 articles.
1. The FAIR Guiding Principles for scientific data management and stewardship
2. Editorial: biological ontologies and semantic biology
3. Precision omics data integration and analysis with interoperable ontologies and their application for COVID-19 research
4. A Survey of Semantic Integration Approaches in Bioinformatics;Messaoudi;Int. J. Comput. Inf. Eng.,2016
5. Mining the Web of Life Sciences Linked Open Data for Mechanism-Based Pharmacovigilance;Kamdar;Proceedings of the WWW’18: Companion Proceedings of the The Web Conference