Exploring topics related to data mining on Wikipedia

Author:

Wang Yanyan,Zhang Jin

Abstract

Purpose Data mining has been a popular research area in the past decades. Many researchers study data-mining theories, methods, applications and trends; however, there are very few studies on data-mining-related topics in social media. This paper aims to explore the topics related to data mining based on the data collected from Wikipedia. Design/methodology/approach In total, 402 data-mining-related articles were obtained from Wikipedia. These articles were manually classified into several categories by the coding method. Each category formed an article-term matrix. These matrices were analysed and visualized by the self-organizing map approach. Several clusters were observed in each category. Finally, the topics of these clusters were extracted by content analysis. Findings The articles obtained were classified into six categories: applications, foundation and concepts, methodologies, organizations, related fields and topics and technology support. Business, biology and security were the three prominent topics of the applications category. The technologies supporting data mining were software, systems, databases, programming languages and so forth. The general public was more interested in data-mining organizations than the researchers. They also focused on the applications of data mining in business more than in other fields. Originality/value This study will help researchers gain insight into the general public’s perceptions of data mining and discover the gap between the general public and themselves. It will assist researchers in finding new techniques and methods which will potentially provide them with new data-mining methods and research topics.

Publisher

Emerald

Subject

Library and Information Sciences,Computer Science Applications

Reference61 articles.

1. Social media road maps exploring the futures triggered by social media;VTT Tiedotteita-Valtion Teknillinen Tutkimuskeskus,2008

2. Application of data mining: diabetes health care in young and old patients;Journal of King Saud University-Computer and Information Sciences,2013

3. The visual subject analysis of library and information science journals with self-organizing map;Knowledge Organization,2011

4. Motivating and discouraging factors for Wikipedians: the case study of Persian Wikipedia;Library Review,2013

5. Commons-based peer production and virtue;Journal of Political Philosophy,2006

Cited by 3 articles. 订阅此论文施引文献 订阅此论文施引文献,注册后可以免费订阅5篇论文的施引文献,订阅后可以查看论文全部施引文献

1. The impact of big data on research methods in information science;Data and Information Management;2023-06

2. Data Mining and Machine Learning Approaches and Technologies for Diagnosing Diabetes in Women;Big Data and Networks Technologies;2019-07-18

3. Clustering in the presence of side information: a non-linear approach;International Journal of Intelligent Computing and Cybernetics;2019-06-10

同舟云学术

1.学者识别学者识别

2.学术分析学术分析

3.人才评估人才评估

"同舟云学术"是以全球学者为主线,采集、加工和组织学术论文而形成的新型学术文献查询和分析系统,可以对全球学者进行文献检索和人才价值评估。用户可以通过关注某些学科领域的顶尖人物而持续追踪该领域的学科进展和研究前沿。经过近期的数据扩容,当前同舟云学术共收录了国内外主流学术期刊6万余种,收集的期刊论文及会议论文总量共计约1.5亿篇,并以每天添加12000余篇中外论文的速度递增。我们也可以为用户提供个性化、定制化的学者数据。欢迎来电咨询!咨询电话:010-8811{复制后删除}0370

www.globalauthorid.com

TOP

Copyright © 2019-2024 北京同舟云网络信息技术有限公司
京公网安备11010802033243号  京ICP备18003416号-3