Fine-grained semantic type discovery for heterogeneous sources using clustering-Reference-Cited by-同舟云学术

Fine-grained semantic type discovery for heterogeneous sources using clustering

Published:2022-05-17 Issue:2 Volume:32 Page:305-324
ISSN:1066-8888
Container-title:The VLDB Journal
language:en
Short-container-title:The VLDB Journal

Author:

Piai Federico^ORCID,Atzeni Paolo^ORCID,Merialdo Paolo^ORCID,Srivastava Divesh^ORCID

Abstract

AbstractWe focus on the key task of semantic type discovery over a set of heterogeneous sources, an important data preparation task. We consider the challenging setting of multiple Web data sources in a vertical domain, which present sparsity of data and a high degree of heterogeneity, even internally within each individual source. We assume each source provides a collection of entity specifications, i.e. entity descriptions, each expressed as a set of attribute name-value pairs. Semantic type discovery aims at clustering individual attribute name-value pairs that represent the same semantic concept. We take advantage of the opportunities arising from the redundancy of information across such sources and propose the iterative RaF-STD solution, which consists of three key steps: (i) a Bayesian model analysis of overlapping information across sources to match the most locally homogeneous attributes; (ii) a tagging approach, inspired by NLP techniques, to create (virtual) homogeneous attributes from portions of heterogeneous attribute values; and (iii) a novel use of classical techniques based on matching of attribute names and domains. Empirical evaluation on the DI2KG and WDC benchmarks demonstrates the superiority of RaF-STD over alternative approaches adapted from the literature.

Publisher

Springer Science and Business Media LLC

Subject

Hardware and Architecture,Information Systems

Link

https://link.springer.com/content/pdf/10.1007/s00778-022-00743-3.pdf

Reference46 articles.

1. Abedjan, Z., Chu, X., Deng, D., Fernandez, R.C., Ilyas, I.F., Ouzzani, M., Papotti, P., Stonebraker, M., Tang, N.: Detecting data errors: Where are we and what needs to be done? PVLDB 9(12), 993–1004 (2016)

2. Aumueller, D., Do, H.H., Massmann, S., Rahm, E.: Schema and ontology matching with coma++. In: Proceedings of the 2005 ACM SIGMOD international conference on Management of data, pp. 906–908 (2005)

3. Banerjee, A., Krumpelman, C., Ghosh, J., Basu, S., Mooney, R.J.: Model-based overlapping clustering. In: Proceedings of the eleventh ACM SIGKDD international conference on Knowledge discovery in data mining, pp. 532–537 (2005)

4. Barbosa, L., Crescenzi, V., Dong, X.L., Merialdo, P., Piai, F., Qiu, D., Shen, Y., Srivastava, D.: Big data integration for product specifications. IEEE Data Eng. Bull. 41(2), 71–81 (2018)

5. Bellahsene, Z., Bonifati, A., Rahm, E.: Schema Matching and Mapping. Springer Science & Business Media, Berlin (2011)

Cited by 2 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Exploring Relationships Between Data in Enterprise Information Systems by Analysis of Log Contents;Lecture Notes in Business Information Processing;2024

2. Dataset Discovery and Exploration: A Survey;ACM Computing Surveys;2023-11-09