Assessing word commonness-Reference-Cited by-同舟云学术

Assessing word commonness

Published:2022-11-25 Issue: Volume: Page:
ISSN:1384-6655
Container-title:International Journal of Corpus Linguistics
language:en
Short-container-title:IJCL

Author:

Ekeland Paulsen Mikkel¹^ORCID

Affiliation:

1. University of Bergen

Abstract

Abstract The article investigates the two main corpus indicators of word commonness, frequency and dispersion, through a cross-validation analysis of frequency and four dispersion measures (‘Range’, ‘Chi-squared’, ‘Deviation of Proportions’ and ‘Juilland’s D’). The approach provides an estimation of the capacity of the named measures to predict the distribution of corpus items in an extracted language sample. Based on a dataset of 273 Norwegian compounds, the results show that especially Deviation of Proportions is a robust measure of dispersion that can be used in conjunction with frequency to substantiate assertions of word commonness based on corpus data. In addition, dispersion measures do not only reflect what sort of distribution the frequency statistic is generated from, but also how reliable the frequency estimation in the corpus sample is in terms of giving an accurate representation of frequency in the language variety that the corpus is sampled from.

Publisher

John Benjamins Publishing Company

Subject

Linguistics and Language,Language and Linguistics

Link

http://www.jbe-platform.com/deliver/fulltext/10.1075/ijcl.21037.eke/ijcl.21037.eke.pdf

Reference17 articles.

1. Analyzing Linguistic Data

2. The Utility of Item-Level Analyses in Model Evaluation: A Reply to Seidenberg and Plaut

3. On the (non)utility of Juilland’s D to measure lexical dispersion in large corpora

Cited by 1 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Wheat or Chaff? A Compound Selection Model Based on Look-Up Data;International Journal Of Lexicography;2023-05-30