Affiliation:
1. Research Group in Computational Linguistics, University of Wolverhampton, UK
Abstract
Abstract
There has not been any research that provides an evaluation of the linguistic features extracted from the matn (text) of a Hadith. Moreover, none of the fairly large corpora are publicly available as a benchmark corpus for Hadith authenticity, and there is a need to build a ‘gold standard’ corpus for good practices in Hadith authentication. We write a scraper in Python programming language and collect a corpus of 3,651 authentic prophetic traditions and 3,593 fake ones. We process the corpora with morphological segmentation and perform extensive experimental studies using a variety of machine learning algorithms, mainly through automatic machine learning, to distinguish between these two categories. With a feature set including words, morphological segments, characters, top N words, top N segments, function words, and several vocabulary richness features, we analyze the results in terms of both prediction and interpretability to explain which features are more characteristic of each class. Many experiments have produced good results and the highest accuracy (i.e. 78.28%) is achieved using word n-grams as features using the Multinomial Naive Bayes classifier. Our extensive experimental studies conclude that, at least for Digital Humanities, feature engineering may still be desirable due to the high interpretability of the features. The corpus and software (scripts) will be made publicly available to other researchers in an effort to promote progress and replicability.
Publisher
Oxford University Press (OUP)
Subject
Computer Science Applications,Linguistics and Language,Language and Linguistics,Information Systems
Reference40 articles.
1. Hadith classification using machine learning techniques according to its reliability;Abdelaal;Science and Technology,2019
2. Spoken and written textual dimensions in English: resolving the contradictory findings;Biber;Language,1986
Cited by
7 articles.
订阅此论文施引文献
订阅此论文施引文献,注册后可以免费订阅5篇论文的施引文献,订阅后可以查看论文全部施引文献