Open data integration-Reference-Cited by-同舟云学术

Open data integration

Published:2018-08 Issue:12 Volume:11 Page:2130-2139
ISSN:2150-8097
Container-title:Proceedings of the VLDB Endowment
language:en
Short-container-title:Proc. VLDB Endow.

Author:

Miller Renée J.¹

Affiliation:

1. Northeastern University

Abstract

Open data plays a major role in supporting both governmental and organizational transparency. Many organizations are adopting Open Data Principles promising to make their open data complete, primary, and timely. These properties make this data tremendously valuable to data scientists. However, scientists generally do not have a priori knowledge about what data is available (its schema or content). Nevertheless, they want to be able to use open data and integrate it with other public or private data they are studying. Traditionally, data integration is done using a framework called query discovery where the main task is to discover a query (or transformation) that translates data from one form into another. The goal is to find the right operators to join, nest, group, link, and twist data into a desired form. We introduce a new paradigm for thinking about integration where the focus is on data discovery, but highly efficient internet-scale discovery that is driven by data analysis needs. We describe a research agenda and recent progress in developing scalable data-analysis or query-aware data discovery algorithms that provide high recall and accuracy over massive data repositories.

Publisher

VLDB Endowment

Subject

General Earth and Planetary Sciences,Water Science and Technology,Geography, Planning and Development

Link

https://dl.acm.org/doi/pdf/10.14778/3229863.3240491

Cited by 59 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. A Large Scale Test Corpus for Semantic Table Search;Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval;2024-07-10

2. Meta-Dataset Creation Method for Classification and Search of Open Datasets;The Journal of Korean Institute of Information Technology;2024-04-30

3. Fast Shapley Value Computation in Data Assemblage Tasks as Cooperative Simple Games;Proceedings of the ACM on Management of Data;2024-03-12

4. Self-supervised data lakes discovery through unsupervised metadata-driven weighted similarity;Information Sciences;2024-03

5. Analytic Processing in Data Lakes: A Semantic Query-Driven Discovery Approach;Information Systems Frontiers;2024-02-14