SimLex-999: Evaluating Semantic Models With (Genuine) Similarity Estimation-Reference-Cited by-同舟云学术

SimLex-999: Evaluating Semantic Models With (Genuine) Similarity Estimation

Published:2015-12 Issue:4 Volume:41 Page:665-695
ISSN:0891-2017
Container-title:Computational Linguistics
language:en
Short-container-title:Computational Linguistics

Author:

Hill Felix¹,Reichart Roi²,Korhonen Anna¹

Affiliation:

1. University of Cambridge

2. Technion, Israel Institute of Technology

Abstract

We present SimLex-999, a gold standard resource for evaluating distributional semantic models that improves on existing resources in several important ways. First, in contrast to gold standards such as WordSim-353 and MEN, it explicitly quantifies similarity rather than association or relatedness so that pairs of entities that are associated but not actually similar (Freud, psychology) have a low rating. We show that, via this focus on similarity, SimLex-999 incentivizes the development of models with a different, and arguably wider, range of applications than those which reflect conceptual association. Second, SimLex-999 contains a range of concrete and abstract adjective, noun, and verb pairs, together with an independent rating of concreteness and (free) association strength for each pair. This diversity enables fine-grained analyses of the performance of models on concepts of different types, and consequently greater insight into how architectures can be improved. Further, unlike existing gold standard evaluations, for which automatic approaches have reached or surpassed the inter-annotator agreement ceiling, state-of-the-art models perform well below this ceiling on SimLex-999. There is therefore plenty of scope for SimLex-999 to quantify future improvements to distributional semantic models, guiding the development of the next generation of representation-learning architectures.

Publisher

MIT Press - Journals

Subject

Artificial Intelligence,Computer Science Applications,Linguistics and Language,Language and Linguistics

Link

https://www.mitpressjournals.org/doi/pdf/10.1162/COLI_a_00237

Reference67 articles.

1. A study on similarity and relatedness using distributional and WordNet-based approaches

2. Integrating experiential and distributional data to learn semantic representations.

3. Bansal, Mohit, Kevin Gimpel, and Karen Livescu. 2014. Tailoring continuous word representations for dependency parsing. In Proceedings of ACL, Baltimore, MD.

4. Baroni, Marco, Georgiana Dinu, and Germán Kruszewski. 2014. Don't count, predict! A systematic comparison of context-counting vs. context-predicting semantic vectors. In Proceedings of ACL, Baltimore, MD.

5. Distributional Memory: A General Framework for Corpus-Based Semantics

Cited by 404 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Improving semantic similarity computation via subgraph feature fusion based on semantic awareness;Engineering Applications of Artificial Intelligence;2024-10

2. Why concepts are (probably) vectors;Trends in Cognitive Sciences;2024-09

3. PrimeNet: A Framework for Commonsense Knowledge Representation and Reasoning Based on Conceptual Primitives;Cognitive Computation;2024-08-30

4. Effect of dimension size and window size on word embedding in classification tasks;2024-07-08

5. Flexible margins and multiple samples learning to enhance lexical semantic similarity;Engineering Applications of Artificial Intelligence;2024-07