Data Lake Governance: Towards a Systemic and Natural Ecosystem Analogy-Reference-Cited by-同舟云学术

Data Lake Governance: Towards a Systemic and Natural Ecosystem Analogy

Published:2020-07-27 Issue:8 Volume:12 Page:126
ISSN:1999-5903
Container-title:Future Internet
language:en
Short-container-title:Future Internet

Author:

Derakhshannia Marzieh^ORCID,Gervet Carmen^ORCID,Hajj-Hassan Hicham^ORCID,Laurent Anne^ORCID,Martin Arnaud^ORCID

Abstract

The realm of big data has brought new venues for knowledge acquisition, but also major challenges including data interoperability and effective management. The great volume of miscellaneous data renders the generation of new knowledge a complex data analysis process. Presently, big data technologies provide multiple solutions and tools towards the semantic analysis of heterogeneous data, including their accessibility and reusability. However, in addition to learning from data, we are faced with the issue of data storage and management in a cost-effective and reliable manner. This is the core topic of this paper. A data lake, inspired by the natural lake, is a centralized data repository that stores all kinds of data in any format and structure. This allows any type of data to be ingested into the data lake without any restriction or normalization. This could lead to a critical problem known as data swamp, which can contain invalid or incoherent data that adds no values for further knowledge acquisition. To deal with the potential avalanche of data, some legislation is required to turn such heterogeneous datasets into manageable data. In this article, we address this problem and propose some solutions concerning innovative methods, derived from a multidisciplinary science perspective to manage data lake. The proposed methods imitate the supply chain management and natural lake principles with an emphasis on the importance of the data life cycle, to implement responsible data governance for the data lake.

Publisher

MDPI AG

Subject

Computer Networks and Communications

Link

https://www.mdpi.com/1999-5903/12/8/126/pdf

Reference67 articles.

1. The FAIR Guiding Principles for scientific data management and stewardship

2. The next information architecture evolution

3. Data lakes: Purposes, practices, patterns, and platforms;Russom,2017

4. Data Lake: A New Ideology in Big Data Era https://www.itm-conferences.org/articles/itmconf/pdf/2018/02/itmconf_wcsn2018_03025.pdf

Cited by 10 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Security and Ownership in User-Defined Data Meshes;Algorithms;2024-04-22

2. Design and Efficacy of a Data Lake Architecture for Multimodal Emotion Feature Extraction in Social Media;IET Software;2024-03-08

3. Data lake governance using IBM-Watson knowledge catalog;Scientific African;2023-09

4. Towards a framework for developing visual analytics in supply chain environments;International Journal of Information Systems and Project Management;2023-04-06

5. DEWA R&D Data Lake: Big Data Platform for Advanced Energy Data Analytics;2023 International Conference on IT Innovation and Knowledge Discovery (ITIKD);2023-03-08