Abstract
PurposeIn order to estimate the value of semi-automated subject indexing in operative library catalogues, the study aimed to investigate five different automated implementations of an open source software package on a large set of Swedish union catalogue metadata records, with Dewey Decimal Classification (DDC) as the target classification system. It also aimed to contribute to the body of research on aboutness and related challenges in automated subject indexing and evaluation.Design/methodology/approachOn a sample of over 230,000 records with close to 12,000 distinct DDC classes, an open source tool Annif, developed by the National Library of Finland, was applied in the following implementations: lexical algorithm, support vector classifier, fastText, Omikuji Bonsai and an ensemble approach combing the former four. A qualitative study involving two senior catalogue librarians and three students of library and information studies was also conducted to investigate the value and inter-rater agreement of automatically assigned classes, on a sample of 60 records.FindingsThe best results were achieved using the ensemble approach that achieved 66.82% accuracy on the three-digit DDC classification task. The qualitative study confirmed earlier studies reporting low inter-rater agreement but also pointed to the potential value of automatically assigned classes as additional access points in information retrieval.Originality/valueThe paper presents an extensive study of automated classification in an operative library catalogue, accompanied by a qualitative study of automated classes. It demonstrates the value of applying semi-automated indexing in operative information retrieval systems.
Reference36 articles.
1. The nature of indexing: how humans and machines analyze messages and texts for retrieval. Part I: research, and the nature of human indexing;Information Processing and Management,2001
2. Making a science of model search: hyperparameter optimization in hundreds of dimensions for vision architectures,2013
3. Why build Dewey numbers? The remediation of the Dewey decimal classification system;Nordlit,2012
4. Taming pretrained transformers for extreme multi-label text classification,2020
5. Conradi, E. (2017), “DDC and automatic classification”, available at: https://edug.pansoft.de/tiki-index.php?page=DDC+and+automatic+classification
Cited by
1 articles.
订阅此论文施引文献
订阅此论文施引文献,注册后可以免费订阅5篇论文的施引文献,订阅后可以查看论文全部施引文献