Author:
Yucel Ahmet,Caglar Musa,Ahady Dolatsara Hamidreza,George Benjamin,Dag Ali
Abstract
Purpose
Machine learning algorithms are useful to effectively analyse, and therefore automatically classify online reviews. The purpose of this paper is to demonstrate a novel text-mining framework and its potential for use in the classification of unstructured hotel reviews.
Design/methodology/approach
Well-known data mining methods (i.e. boosted decision trees (BDT), classification and regression trees (C&RT) and random forests (RF)) in conjunction with incorporating five-fold cross-validation are used to predict the star rating of the hotel reviews. To achieve this goal, extracted features are used to create a composite variable (CV) to deploy into machine learning algorithms as the main feature (variable) during the learning process.
Findings
BDT outperformed the other alternatives in the exact accuracy rate (EAR) and multi-class accuracy rate (MCAR) by reaching the accuracy rates of 0.66 and 0.899, respectively. Moreover, phrases such as “clean”, “friendly”, “nice”, “perfect” and “love” are shown to be associated with four and five stars, whereas, phrases such as “horrible”, “never”, “terrible” and “worst” are shown to be associated with one and two-star hotels, as it would be the intuitive expectation.
Originality/value
To the best of the knowledge, there is no study in the existent literature, which synthesizes the knowledge obtained from individual features and uses them to create a single composite variable that is powerful enough to predict the star rates of the user-generated reviews. This study believes that the proposed method also provides policymakers with a unique window in the thoughts and opinions of individual users, which may be used to augment the current decision-making process.
Subject
Management Science and Operations Research,Strategy and Management,General Decision Sciences
Reference46 articles.
1. Assessing text mining alogrithm outcomes;Journal of Business Analytics,2020
2. Data mining for credit card fraud: a comparative study;Decision Support Systems,2011
3. A study of opinion mining and visualization of hotel reviews,2012
4. A machine learning approach to sentiment analysis in multilingual web texts;Information Retrieval,2009
Cited by
2 articles.
订阅此论文施引文献
订阅此论文施引文献,注册后可以免费订阅5篇论文的施引文献,订阅后可以查看论文全部施引文献