Semantic concept model using Wikipedia semantic features-Reference-Cited by-同舟云学术

Semantic concept model using Wikipedia semantic features

Published:2017-05-02 Issue:4 Volume:44 Page:526-551
ISSN:0165-5515
Container-title:Journal of Information Science
language:en
Short-container-title:Journal of Information Science

Author:

Saif Abdulgabbar¹,Omar Nazlia²,Ab Aziz Mohd Juzaiddin²,Zainodin Ummi Zakiah²,Salim Naomie³

Affiliation:

1. Computer-Math Department, Faculty of Applied Sciences, Thamar University, Yemen; Center for Artificial Intelligence Technology, Faculty of Information Science & Technology, Universiti Kebangsaan Malaysia, Malaysia

2. Center for Artificial Intelligence Technology, Faculty of Information Science & Technology, Universiti Kebangsaan Malaysia, Malaysia

3. Faculty of Computer Science & Information System, Universiti Teknologi Malaysia, Malaysia

Abstract

Wikipedia has become a high coverage knowledge source which has been used in many research areas such as natural language processing, text mining and information retrieval. Several methods have been introduced for extracting explicit or implicit relations from Wikipedia to represent semantics of concepts/words. However, the main challenge in semantic representation is how to incorporate different types of semantic relations to capture more semantic evidences of the associations of concepts. In this article, we propose a semantic concept model that incorporates different types of semantic features extracting from Wikipedia. For each concept that corresponds to an article, four semantic features are introduced: template links, categories, salient concepts and topics. The proposed model is based on the probability distributions that are defined for these semantic features of a Wikipedia concept. The template links and categories are the document-level features which are directly extracted from the structured information included in the article. On the other hand, the salient concepts and topics are corpus-level features which are extracted to capture implicit relations among concepts. For the salient concepts feature, the distributional-based method is utilised on the hypertext corpus to extract this feature for each Wikipedia concept. Then, the probability product kernel is used to improve the weight of each concept in this feature. For the topic feature, the Labelled latent Dirichlet allocation is adapted on the supervised multi-label of Wikipedia to train the probabilistic model of this feature. Finally, we used the linear interpolation for incorporating these semantic features into the probabilistic model to estimate the semantic relation probability of the specific concept over Wikipedia articles. The proposed model is evaluated on 12 benchmark datasets in three natural language processing tasks: measuring the semantic relatedness of concepts/words in general and in the biomedical domain, semantic textual relatedness measurement and measuring the semantic compositionality of noun compounds. The model is also compared with five methods that depends on separate semantic features in Wikipedia. Experimental results show that the proposed model achieves promising results in three tasks and outperforms the baseline methods in most of the evaluation datasets. This implies that incorporation of explicit and implicit semantic features is useful for representing semantics of concepts in Wikipedia.

Publisher

SAGE Publications

Subject

Library and Information Sciences,Information Systems

Link

http://journals.sagepub.com/doi/pdf/10.1177/0165551517706231

Reference105 articles.

1. WordNet

2. BabelNet: The automatic construction, evaluation and application of a wide-coverage multilingual semantic network

Cited by 5 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. The Kadu Lexicon local wisdom of geographic’s toponymic at Pandeglang Regency, Banten Province;IOP Conference Series: Earth and Environmental Science;2022-11-01

2. A Novel Deep Auto-Encoder Based Linguistics Clustering Model for Social Text;ACM Transactions on Asian and Low-Resource Language Information Processing;2022-04-29

3. Co-occurrence graph-based context adaptation: a new unsupervised approach to word sense disambiguation;Digital Scholarship in the Humanities;2020-04-13

4. Assessing semantic similarity between concepts: A weighted‐feature‐based approach;Concurrency and Computation: Practice and Experience;2020-01-03

5. SRL-ESA-TextSum: A text summarization approach based on semantic role labeling and explicit semantic analysis;Information Processing & Management;2019-07