‘Everything is data’: towards one big data ecosystem using multiple sources of data on higher education in Indonesia-Reference-Cited by-同舟云学术

‘Everything is data’: towards one big data ecosystem using multiple sources of data on higher education in Indonesia

Published:2022-07-14 Issue:1 Volume:9 Page:
ISSN:2196-1115
Container-title:Journal of Big Data
language:en
Short-container-title:J Big Data

Author:

Yunita Ariana,Santoso Harry B.,Hasibuan Zainal A.

Abstract

AbstractBig data is increasingly being promoted as a game changer for the future of science, as the volume of data has exploded in recent years. Big data characterized, among others, the data comes from multiple sources, multi-format, comply to 5-V’s in nature (value, volume, velocity, variety, and veracity). Big data also constitutes structured data, semi-structured data, and unstructured-data. These characteristics of big data formed “big data ecosystem” that have various active nodes involved. Regardless such complex characteristics of big data, the studies show that there exists inherent structure that can be very useful to provide meaningful solutions for various problems. One of the problems is anticipating proper action to students’ achievement. It is common practice that lecturer treat his/her class with “one-size-fits-all” policy and strategy. Whilst, the degree of students’ understanding, due to several factors, may not the same. Furthermore, it is often too late to take action to rescue the student’s achievement in trouble. This study attempted to gather all possible features involved from multiple data sources: national education databases, reports, webpages and so forth. The multiple data sources comprise data on undergraduate students from 13 provinces in Indonesia, including students’ academic histories, demographic profiles and socioeconomic backgrounds and institutional information (i.e. level of accreditation, programmes of study, type of university, geographical location). Gathered data is furthermore preprocessed using various techniques to overcome missing value, data categorisation, data consistency, data quality assurance, to produce relatively clean and sound big dataset. Principal component analysis (PCA) is employed in order to reduce dimensions of big dataset and furthermore use K-Means methods to reveal clusters (inherent structure) that may occur in that big dataset. There are 7 clusters suggested by K-Means analysis: 1. very low-risk students, 2. low-risk students, 3. moderate-risk students, 4. fluctuating-risk students, 5. high risk students, 6. very high-risk students and, 7. fail students. Among the clusters unreveal, (1) a gap between public universities and private universities across the three regions in Indonesia, (2) a gap between STEM and non-STEM programmes of study, (3) a gap between rural versus urban, (4) a gap of accreditation status, (5) a gap of quality human resources distribution, etc. Further study, we will use the characteristics of each cluster to predict students’ achievement based on students’ profiles, and provide solutions and interventions strategies for students to improve their likely success.

Funder

Universitas Indonesia

Publisher

Springer Science and Business Media LLC

Subject

Information Systems and Management,Computer Networks and Communications,Hardware and Architecture,Information Systems

Link

https://link.springer.com/content/pdf/10.1186/s40537-022-00639-7.pdf

Reference44 articles.

1. Rydning DR-JG-J, others. The digitization of the world from edge to core. Fram. Int. Data Corp. 2018 [cited 2021 Dec 25]. p. 16. https://www.seagate.com/files/www-content/our-story/trends/files/idc-seagate-dataage-whitepaper.pdf

2. Wu C, Buyya R, Ramamohanarao K. Big data analytics = machine learning + cloud computing. In: Buyya R, Calheiros RN, Dastjerdi AV, editors. Big Data Princ Paradig. Morgan Kaufmann; 2016. p. 1–13.

3. Raut RD, Mangla SK, Narwane VS, Dora M, Liu M. Big Data Analytics as a mediator in Lean, Agile, Resilient, and Green (LARG) practices effects on sustainable supply chains. Transp Res Part E Logist Transp Rev. 2021;145:102170. https://doi.org/10.1016/j.tre.2020.102170.

4. Anshari M, Almunawar MN, Lim SA, Al-Mudimigh A. Customer relationship management and big data enabled: Personalization & customization of services. Appl Comput Informatics. 2019;15:94–101. https://doi.org/10.1016/j.aci.2018.05.004.

5. Aloqool A, Alharafsheh M, Abdellatif H, Alghasawneh LAS, Al-Gasawneh JA. The mediating role of customer relationship management between e-supply chain management and competitive advantage. Int J Data Netw Sci. 2022;6:263–72. https://doi.org/10.5267/J.IJDNS.2021.9.002.

Cited by 9 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Teacher Empowerment in Creative Economy Education: A Case Study at Sd Ta'mirul Islam Surakarta Indonesia;Revista de Gestão Social e Ambiental;2024-04-03

2. Understanding the development of public data ecosystems: from a conceptual model to a six-generation model of the evolution of public data ecosystems;SSRN Electronic Journal;2024

3. The Application of Big Data Technology in Monitoring and Analyzing the Operation of Economic Policies;Learning and Analytics in Intelligent Systems;2024

4. Development of an Evidence-Based Tool to Assess the Relative Vulnerability of Different Communities to Tuberculosis;Kesmas: Jurnal Kesehatan Masyarakat Nasional;2023-11-29

5. Dimensionality reduction model based on integer planning for the analysis of key indicators affecting life expectancy;Journal of Data and Information Science;2023-11-01