Imparting interpretability to word embeddings while preserving semantic structure-Reference-Cited by-同舟云学术

Imparting interpretability to word embeddings while preserving semantic structure

Published:2020-06-09 Issue:6 Volume:27 Page:721-746
ISSN:1351-3249
Container-title:Natural Language Engineering
language:en
Short-container-title:Nat. Lang. Eng.

Author:

Şenel Lütfi Kerem,Utlu İhsan,Şahinuç Furkan,Ozaktas Haldun M.,Koç Aykut

Abstract

AbstractAs a ubiquitous method in natural language processing, word embeddings are extensively employed to map semantic properties of words into a dense vector representation. They capture semantic and syntactic relations among words, but the vectors corresponding to the words are only meaningful relative to each other. Neither the vector nor its dimensions have any absolute, interpretable meaning. We introduce an additive modification to the objective function of the embedding learning algorithm that encourages the embedding vectors of words that are semantically related to a predefined concept to take larger values along a specified dimension, while leaving the original semantic learning mechanism mostly unaffected. In other words, we align words that are already determined to be related, along predefined concepts. Therefore, we impart interpretability to the word embedding by assigning meaning to its vector dimensions. The predefined concepts are derived from an external lexical resource, which in this paper is chosen as Roget’s Thesaurus. We observe that alignment along the chosen concepts is not limited to words in the thesaurus and extends to other related words as well. We quantify the extent of interpretability and assignment of meaning from our experimental results. Manual human evaluation results have also been presented to further verify that the proposed method increases interpretability. We also demonstrate the preservation of semantic coherence of the resulting vector space using word-analogy/word-similarity tests and a downstream task. These tests show that the interpretability-imparted word embeddings that are obtained by the proposed framework do not sacrifice performances in common benchmark tests.

Publisher

Cambridge University Press (CUP)

Subject

Artificial Intelligence,Linguistics and Language,Language and Linguistics,Software

Reference57 articles.

1. Socher, R. , Perelygin, A. , Wu, J. , Chuang, J. , Manning, C.D. , Ng, A.Y. and Potts, C. (2013). Recursive deep models for semantic compositionality over a sentiment treebank. In Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing (EMNLP). Seattle, WA, USA: Association for Computational Linguistics, pp. 1631–1642.

2. From Word To Sense Embeddings: A Survey on Vector Representations of Meaning

3. Embedding a Semantic Network in a Word Space

4. RC-NET

Cited by 4 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. An Interpretable Alternative to Neural Representation Learning for Rating Prediction - Transparent Latent Class Modeling of User Reviews;2024 International Joint Conference on Neural Networks (IJCNN);2024-06-30

2. An incremental clustering algorithm based on semantic concepts;Knowledge and Information Systems;2024-02-15

3. Learning interpretable word embeddings via bidirectional alignment of dimensions with semantic concepts;Information Processing & Management;2022-05

4. Gender bias in legal corpora and debiasing it;Natural Language Engineering;2022-03-30