Abstract
This overview traced the evolution of data management, transitioning from traditional ETL processes to addressing contemporary challenges in Big Data, with a particular emphasis on data ingestion and schema matching. It explored the classification of data ingestion into batch, real-time, and hybrid processing, underscoring the challenges associated with data quality and heterogeneity. Central to the discussion was the role of schema mapping in data alignment, proving indispensable for linking diverse data sources. Recent advancements, notably the adoption of machine learning techniques, were significantly reshaping the landscape. The paper also addressed current challenges, including the integration of new technologies and the necessity for effective schema matching solutions, highlighting the continuously evolving nature of schema matching in the context of Big Data
Publisher
Salud, Ciencia y Tecnologia
Reference43 articles.
1. Souibgui M, Atigui F, Zammali S, Cherfi S, Yahia SB. Data quality in ETL process: A preliminary study. Procedia Computer Science [Internet]. 2019;159. Available from: https://doi.org/10.1016/j.procs.2019.09.223
2. Informatica [Internet]. [cited 2023 Oct 18]. What Is Data Ingestion? Available from: https://www.informatica.com/resources/articles/what-is-data-ingestion.html
3. Alserafi A. Dataset Proximity Mining for Supporting Schema Matching and Data Lake Governance [PhD Thesis]. Universitat Politècnica de Catalunya, BarcelonaTech; 2021.
4. Meehan J, Tatbul N, Aslantas C, Zdonik S. Data Ingestion for the Connected World. In: CIDR’17. 2017.
5. Hoseini S, Ali A, Shaker H, Quix C. SEDAR: A Semantic Data Reservoir for Heterogeneous Datasets. In: 32nd ACM International Conference on Information and Knowledge Management [Internet]. ACM; 2023. p. 5056–60. Available from: https://doi.org/10.1145/3583780.3614753