Word Sense Disambiguation Using Cosine Similarity Collaborates with Word2vec and WordNet-Reference-Cited by-同舟云学术

Word Sense Disambiguation Using Cosine Similarity Collaborates with Word2vec and WordNet

Published:2019-05-12 Issue:5 Volume:11 Page:114
ISSN:1999-5903
Container-title:Future Internet
language:en
Short-container-title:Future Internet

Author:

Orkphol Korawit^ORCID,Yang Wu

Abstract

Words have different meanings (i.e., senses) depending on the context. Disambiguating the correct sense is important and a challenging task for natural language processing. An intuitive way is to select the highest similarity between the context and sense definitions provided by a large lexical database of English, WordNet. In this database, nouns, verbs, adjectives, and adverbs are grouped into sets of cognitive synonyms interlinked through conceptual semantics and lexicon relations. Traditional unsupervised approaches compute similarity by counting overlapping words between the context and sense definitions which must match exactly. Similarity should compute based on how words are related rather than overlapping by representing the context and sense definitions on a vector space model and analyzing distributional semantic relationships among them using latent semantic analysis (LSA). When a corpus of text becomes more massive, LSA consumes much more memory and is not flexible to train a huge corpus of text. A word-embedding approach has an advantage in this issue. Word2vec is a popular word-embedding approach that represents words on a fix-sized vector space model through either the skip-gram or continuous bag-of-words (CBOW) model. Word2vec is also effectively capturing semantic and syntactic word similarities from a huge corpus of text better than LSA. Our method used Word2vec to construct a context sentence vector, and sense definition vectors then give each word sense a score using cosine similarity to compute the similarity between those sentence vectors. The sense definition also expanded with sense relations retrieved from WordNet. If the score is not higher than a specific threshold, the score will be combined with the probability of that sense distribution learned from a large sense-tagged corpus, SEMCOR. The possible answer senses can be obtained from high scores. Our method shows that the result (50.9% or 48.7% without the probability of sense distribution) is higher than the baselines (i.e., original, simplified, adapted and LSA Lesk) and outperforms many unsupervised systems participating in the SENSEVAL-3 English lexical sample task.

Funder

National Key Research and Development Plan

Publisher

MDPI AG

Subject

Computer Networks and Communications

Link

https://www.mdpi.com/1999-5903/11/5/114/pdf

Reference53 articles.

Cited by 55 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Scientific paper recommender system using deep learning and link prediction in citation network;Heliyon;2024-07

2. Content-based medical image retrieval using deep learning-based features and hybrid meta-heuristic optimization;Biomedical Signal Processing and Control;2024-06

3. How Contentious Terms About People and Cultures are Used in Linked Open Data;Proceedings of the ACM Web Conference 2024;2024-05-13

4. High-performance computing in healthcare: An automatic literature analysis perspective;Journal of Big Data;2024-05-02

5. H-sim: uma função de similaridade híbrida para identificação de correspondência de produtos;Revista Brasileira de Computação Aplicada;2024-05-01