Overview on Data Ingestion and Schema Matching-Reference-Cited by-同舟云学术

Overview on Data Ingestion and Schema Matching

Published:2024-08-02 Issue: Volume:3 Page:219
ISSN:2953-4917
Container-title:Data and Metadata
language:
Short-container-title:Data and Metadata

Author:

El Haddadi Oumaima^ORCID,Chevalier Max^ORCID,Dousset Bernard,El Allaoui Ahmad^ORCID,El Haddadi Anass^ORCID,Teste Olivier^ORCID

Abstract

This overview traced the evolution of data management, transitioning from traditional ETL processes to addressing contemporary challenges in Big Data, with a particular emphasis on data ingestion and schema matching. It explored the classification of data ingestion into batch, real-time, and hybrid processing, underscoring the challenges associated with data quality and heterogeneity. Central to the discussion was the role of schema mapping in data alignment, proving indispensable for linking diverse data sources. Recent advancements, notably the adoption of machine learning techniques, were significantly reshaping the landscape. The paper also addressed current challenges, including the integration of new technologies and the necessity for effective schema matching solutions, highlighting the continuously evolving nature of schema matching in the context of Big Data

Publisher

Salud, Ciencia y Tecnologia

Reference43 articles.

1. Souibgui M, Atigui F, Zammali S, Cherfi S, Yahia SB. Data quality in ETL process: A preliminary study. Procedia Computer Science [Internet]. 2019;159. Available from: https://doi.org/10.1016/j.procs.2019.09.223

2. Informatica [Internet]. [cited 2023 Oct 18]. What Is Data Ingestion? Available from: https://www.informatica.com/resources/articles/what-is-data-ingestion.html

3. Alserafi A. Dataset Proximity Mining for Supporting Schema Matching and Data Lake Governance [PhD Thesis]. Universitat Politècnica de Catalunya, BarcelonaTech; 2021.

4. Meehan J, Tatbul N, Aslantas C, Zdonik S. Data Ingestion for the Connected World. In: CIDR’17. 2017.

5. Hoseini S, Ali A, Shaker H, Quix C. SEDAR: A Semantic Data Reservoir for Heterogeneous Datasets. In: 32nd ACM International Conference on Information and Knowledge Management [Internet]. ACM; 2023. p. 5056–60. Available from: https://doi.org/10.1145/3583780.3614753