Duplicate bibliographic record detection with an OCR-converted source of information-Reference-Cited by-同舟云学术

Duplicate bibliographic record detection with an OCR-converted source of information

Published:2012-10-15 Issue:2 Volume:39 Page:153-168
ISSN:0165-5515
Container-title:Journal of Information Science
language:en
Short-container-title:Journal of Information Science

Author:

Taniguchi Shoichi¹

Affiliation:

1. School of Library and Information Science, Keio University, Japan

Abstract

Duplicate record detection has been an important issue in the fields of data and records management and various detection methods have been proposed. A new method, which uses an optical character recognition (OCR)-converted source of information for record matching to detect duplicates, is proposed and examined in this paper. First, the design of an experiment for examining the performance of such a duplicate detection method is discussed. The base record set with an OCR-converted title page and its verso were prepared along with two test record sets from different union catalogues, and duplicate records between the base record set and the test sets were manually identified. A duplicate detection system was developed to execute matching (1) between records, (2) between a record and an OCR-converted source of information and (3) using a combination of these. Second, matching performance at the individual data element level is examined. Third, the performance of duplicate record detection based on matching at the element level is examined through rule-based detection and machine learning-based detection. The results of the experiment show the usefulness of incorporating source of information into duplicate detection to a certain extent.

Publisher

SAGE Publications

Subject

Library and Information Sciences,Information Systems

Link

http://journals.sagepub.com/doi/pdf/10.1177/0165551512459923

Reference23 articles.

1. Duplicate detection and record consolidation in large bibliographic databases: the COPAC database experience

2. An expert system for quality control and duplicate detection in bibliographic databases

3. Duplicate detection algorithms of bibliographic descriptions

Cited by 4 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Duplicate Records in WorldCat for 20th-Century American, British, and Canadian Books: A Comparison of Duplication Rates and Causes;Cataloging & Classification Quarterly;2024-02-17

2. Locality sensitive blocking (LSB): A robust blocking technique for data deduplication;Journal of Information Science;2022-09-16

3. A node resistance-based probability model for resolving duplicate named entities;Scientometrics;2020-07-13

4. Partition Aware Duplicate Records Detection (PADRD) Methodology in Big Data - Decision Support Systems;Communications in Computer and Information Science;2018