Identifying Anomalous Data Entries in Repeated Surveys-Reference-Cited by-同舟云学术

Identifying Anomalous Data Entries in Repeated Surveys

Published:2024 Issue: Volume: Page:436-455
ISSN:1680-743X
Container-title:Journal of Data Science
language:en
Short-container-title:

Author:

Sartore Luca^ORCID,Chen Lu^ORCID,van Wart Justin,Dau Andrew^ORCID,Bejleri Valbona^ORCID

Abstract

The presence of outliers in a dataset can substantially bias the results of statistical analyses. In general, micro edits are often performed manually on all records to correct for outliers. A set of constraints and decision rules is used to simplify the editing process. However, agricultural data collected through repeated surveys are characterized by complex relationships that make revision and vetting challenging. Therefore, maintaining high data-quality standards is not sustainable in short timeframes. The United States Department of Agriculture’s (USDA’s) National Agricultural Statistics Service (NASS) has partially automated its editing process to improve the accuracy of final estimates. NASS has investigated several methods to modernize its anomaly detection system because simple decision rules may not detect anomalies that break linear relationships. In this article, a computationally efficient method that identifies format-inconsistent, historical, tail, and relational anomalies at the data-entry level is introduced. Four separate scores (i.e., one for each anomaly type) are computed for all nonmissing values in a dataset. A distribution-free method motivated by the Bienaymé-Chebyshev’s inequality is used for scoring the data entries. Fuzzy logic is then considered for combining four individual scores into one final score to determine the outliers. The performance of the proposed approach is illustrated with an application to NASS survey data.

Publisher

School of Statistics, Renmin University of China

Reference25 articles.

1. Robust estimation of multivariate location and scatter in the presence of cellwise and casewise contamination;Test,2015

2. Propagation of outliers in multivariate data;The Annals of Statistics,2009

3. Considérations à l’appui de la découverte de Laplace sur la loi de probabilité dans la méthode des moindres carrés;Journal de Mathématiques Pures et Appliquées,1867

4. On outlier detection with the Chebyshev type inequalities;Journal of the Belarusian State University. Mathematics and Informatics,2020

5. OpenMP: An industry standard API for shared-memory programming;IEEE Computational Science and Engineering,1998

Cited by 1 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Introduction to the GASP Special Issue;Journal of Data Science;2024