Towards More Accurate Statistical Profiling of Deployed schema.org Microdata-Reference-Cited by-同舟云学术

Towards More Accurate Statistical Profiling of Deployed schema.org Microdata

Published:2016-11-29 Issue:1 Volume:8 Page:1-31
ISSN:1936-1955
Container-title:Journal of Data and Information Quality
language:en
Short-container-title:J. Data and Information Quality

Author:

Meusel Robert¹,Ritze Dominique¹,Paulheim Heiko¹

Affiliation:

1. Research Group Data and Web Science, University of Mannheim, Mannheim, Germany

Abstract

Being promoted by major search engines such as Google, Yahoo!, Bing, and Yandex, Microdata embedded in web pages, especially using schema.org, has become one of the most important markup languages for the Web. However, deployed Microdata is very often not free from errors, which makes it difficult to estimate the data volume and create an accurate data profile. In addition, as the usage of global identifiers is not common, the real number of entities described by this format in the Web is hard to assess. In this article, we discuss how the subsequent application of data cleaning steps, such as duplicate detection and correction of common schema-based errors, leads to a more realistic view on the data, step by step. The cleaning steps applied include both heuristics for fixing errors as well as means to perform duplicate detection and duplicate elimination. Using the Web Data Commons Microdata corpus, we show that applying such quality improvement methods can essentially change the statistical profile of the dataset and lead to different estimates of both the number of entities as well as the class distribution within the data.

Funder

Amazon Web Service Education

Publisher

Association for Computing Machinery (ACM)

Subject

Information Systems and Management,Information Systems

Link

https://dl.acm.org/doi/pdf/10.1145/2992788

Reference47 articles.

1. Profiling and mining RDF data with ProLOD++

2. Reconciling ontologies and the web of data

3. LOD Laundromat: A Uniform Way of Publishing Other People’s Dirty Data

4. Deployment of RDFa, Microdata, and Microformats on the Web – A Quantitative Analysis

5. Linked data-the story so far;Bizer Christian;Sem. Serv. Interoper. Web Appl.: Emerg. Concepts,2009

Cited by 3 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. KnowMore – knowledge base augmentation with structured web markup;Semantic Web;2018-12-28

2. Inferring Missing Categorical Information in Noisy and Sparse Web Markup;Proceedings of the 2018 World Wide Web Conference on World Wide Web - WWW '18;2018

3. Data Integration for Open Data on the Web;Reasoning Web. Semantic Interoperability on the Web;2017