Affiliation:
1. Westat Inc., Rockville, MD, USA
Abstract
During data collection, field interviewers often append notes or comments to a case in open text fields to request updates to case-level data. Processing these comments can improve data quality, but many are non-actionable, and processing remains a costly manual task. This article presents a case study using a novel application of machine learning tools to assist in the evaluation of these comments. Using over 5,000 comments from the Medical Expenditure Panel Survey, we built features that were fed to a machine learning model to predict a grouping category for each comment as previously assigned by data technicians to expedite processing. The model achieved high top-3 accuracy and was incorporated into a production tool for editing. A qualitative evaluation of the tool also provided encouraging results. This application of machine learning tools allowed a small but worthwhile increase in processing efficiency, while maintaining exacting standards for data quality.
Reference26 articles.
1. Athey L., Kennickell A. B. 2005. Managing data quality on the 2004 Survey of Consumer Finances. Paper presented at the Annual Meetings of AAPOR, Miami. http://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.182.149&rep=rep1&type=pdf (accessed August 1, 2020).
2. Bricker J., Moore K., Windle R. 2014. Examining interviewer–respondent interactions in the Survey of Consumer Finances (SCF). In Proceedings of the 2014 Survey Research Methods Section of ASA, 2162–68. http://www.asasrms.org/Proceedings/y2014/files/312089_88725.pdf (accessed August 1, 2020).
3. Questionnaires and Lived Experience: Strategies of Coping With the Quantitative Frame
4. Automatic Coding of Text Answers to Open-Ended Questions: Should You Double Code the Training Data?
Cited by
1 articles.
订阅此论文施引文献
订阅此论文施引文献,注册后可以免费订阅5篇论文的施引文献,订阅后可以查看论文全部施引文献