Predictability of Word Forms (Types) and Lemmas in Linguistic Corpora. A Case Study Based on the Analysis of the CUMBRE Corpus-Reference-Cited by-同舟云学术

Predictability of Word Forms (Types) and Lemmas in Linguistic Corpora. A Case Study Based on the Analysis of the CUMBRE Corpus

Published:1997-01-01 Issue:2 Volume:2 Page:259-280
ISSN:1384-6655
Container-title:International Journal of Corpus Linguistics
language:en
Short-container-title:IJCL

Author:

Sánchez Aquilino¹,Cantos-Gomez Pascual¹

Affiliation:

1. Universidad de Murcia

Abstract

Various research centres and publishing companies all around the world have been developing corpus resources for many years, and there has been a growing awareness throughout the eighties of their importance to linguistic and lexicographic work. To give some idea of scale, the British National Corpus contains 100 million words, and its counterpart for Spanish—compiled by the Spanish Real Academia de la Lengua—will reach 100 million words at first and 200 million words in a second stage. However, little convincing research has been done in the direction of sample size—directly connected to a further topic: representativeness. We shall investigate here a related issue: Is it possible to predict the different word forms and lemmas of a given corpus? And if so, how? A positive answer to this question may contribute to decision making regarding some aspects of representativeness in given fields. We shall attempt further to find a reliable procedure to predict the total number of word forms (types) and lemmas in a specific corpus.

Publisher

John Benjamins Publishing Company

Subject

Linguistics and Language,Language and Linguistics

Link

http://www.jbe-platform.com/deliver/fulltext/ijcl.2.2.06san.pdf

Cited by 16 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. TheCorpus of Contemporary English Legal Decisions, 1950–2021 (CoCELD): A new tool for analysing recent changes in English legal discourse;ICAME Journal;2023-05-01

2. Lemmatization of Inflected Nouns;Language Corpora Annotation and Processing;2021

3. Corpus and Technical TermBank;Utility and Application of Language Corpora;2018-08-14

4. Processing Texts in a Corpus;Utility and Application of Language Corpora;2018-08-14

5. References;English Corpus Linguistics;2002-06-13