Annif: DIY automated subject indexing using multiple algorithms-Reference-Cited by-同舟云学术

Annif: DIY automated subject indexing using multiple algorithms

Published:2019-07-29 Issue:1 Volume:29 Page:1-25
ISSN:2213-056X
Container-title:LIBER Quarterly: The Journal of the Association of European Research Libraries
language:
Short-container-title:LIBER

Author:

Suominen Osma

Abstract

Manually indexing documents for subject-based access is a labour-intensive process. We propose using metadata gathered from bibliographic databases to train algorithms that assist librarians in that work. We have developed Annif, an open source tool and microservice for automated subject indexing. After training it with a subject vocabulary and existing metadata, Annif can be used to assign subject headings for new documents. We have tested Annif with different document collections including scientific papers, old scanned books and contemporary e-books, Q&A pairs from an “ask a librarian” service, Finnish Wikipedia, and the archives of a local newspaper. The results of analysing scientific papers and current books have been reassuring, while other types of documents have proved to be more challenging. The current version is based on a combination of existing natural language processing and machine learning tools. By combining multiple approaches and existing open source algorithms, Annif can build on the strengths of individual algorithms and adapt to different settings. With Annif, we expect to improve subject indexing and classification processes especially for electronic documents as well as collections that otherwise would not be indexed at all.

Publisher

Ligue des Bibliotheques Europeennes de Recherche

Subject

Library and Information Sciences

Cited by 19 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Pametne knjižnice;Knjižnica: revija za področje bibliotekarstva in informacijske znanosti;2023-12-29

2. Machine Learning Applications in Digital Humanities: Designing a Semi-automated Subject Indexing System for a Low-resource Domain;DESIDOC J LIB INF TE;2023

3. Automatic Indexing for Agriculture: Designing a Framework by Deploying Agrovoc, Agris and Annif;Journal of Information and Knowledge;2023-05-13

4. Machine Learning and Bibliographic Data Universe: Assessing Efficacy of Backend Algorithms in Annif through Retrieval Metrics;SRELS Journal of Information Management;2023-03-27

5. Automated Knowledge Organisation: AI/ML-based Subject Indexing System for Libraries;DESIDOC J LIB INF TE;2023