Developing a big data analytics platform using Apache Hadoop Ecosystem for delivering big data services in libraries-Reference-Cited by-同舟云学术

Developing a big data analytics platform using Apache Hadoop Ecosystem for delivering big data services in libraries

Published:2024-02-22 Issue:2 Volume:40 Page:160-186
ISSN:2059-5816
Container-title:Digital Library Perspectives
language:en
Short-container-title:DLP

Author:

Singh Ranjeet Kumar

Abstract

Purpose Although the challenges associated with big data are increasing, the question of the most suitable big data analytics (BDA) platform in libraries is always significant. The purpose of this study is to propose a solution to this problem. Design/methodology/approach The current study identifies relevant literature and provides a review of big data adoption in libraries. It also presents a step-by-step guide for the development of a BDA platform using the Apache Hadoop Ecosystem. To test the system, an analysis of library big data using Apache Pig, which is a tool from the Apache Hadoop Ecosystem, was performed. It establishes the effectiveness of Apache Hadoop Ecosystem as a powerful BDA solution in libraries. Findings It can be inferred from the literature that libraries and librarians have not taken the possibility of big data services in libraries very seriously. Also, the literature suggests that there is no significant effort made to establish any BDA architecture in libraries. This study establishes the Apache Hadoop Ecosystem as a possible solution for delivering BDA services in libraries. Research limitations/implications The present work suggests adapting the idea of providing various big data services in a library by developing a BDA platform, for instance, providing assistance to the researchers in understanding the big data, cleaning and curation of big data by skilled and experienced data managers and providing the infrastructural support to store, process, manage, analyze and visualize the big data. Practical implications The study concludes that Apache Hadoops’ Hadoop Distributed File System and MapReduce components significantly reduce the complexities of big data storage and processing, respectively, and Apache Pig, using Pig Latin scripting language, is very efficient in processing big data and responding to queries with a quick response time. Originality/value According to the study, there are significantly fewer efforts made to analyze big data from libraries. Furthermore, it has been discovered that acceptance of the Apache Hadoop Ecosystem as a solution to big data problems in libraries are not widely discussed in the literature, although Apache Hadoop is regarded as one of the best frameworks for big data handling.

Publisher

Emerald

Reference68 articles.

1. An analysis of academic librarians competencies and skills for implementation of big data analytics in libraries;Data Technologies and Applications,2019

2. Librarian’s perspective for the implementation of big data analytics in libraries on the bases of lean-startup model;Digital Library Perspectives,2019

3. Performance analysis of ECG big data using Apache Hive and Apache Pig,2019

4. Defining big data and measuring its associated trends in the field of information and library management;Library Hi Tech News,2017

5. Big data adoption in academic libraries: a literature review;Library Hi Tech News,2020

Cited by 2 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Investigation of research support services (RSS) in academic libraries of India;Journal of Librarianship and Information Science;2024-04-17

2. Research on Project-based Teaching Reform of Nautical English in Higher Vocational Colleges under the Background of Informatization;Applied Mathematics and Nonlinear Sciences;2024-01-01