LanguageCrawl: a generic tool for building language models upon common Crawl-Reference-Cited by-同舟云学术

LanguageCrawl: a generic tool for building language models upon common Crawl

Published:2021-08-05 Issue:4 Volume:55 Page:1047-1075
ISSN:1574-020X
Container-title:Language Resources and Evaluation
language:en
Short-container-title:Lang Resources & Evaluation

Author:

Roziewski Szymon^ORCID,Kozłowski Marek^ORCID

Abstract

AbstractThe exponential growth of the internet community has resulted in the production of a vast amount of unstructured data, including web pages, blogs and social media. Such a volume consisting of hundreds of billions of words is unlikely to be analyzed by humans. In this work we introduce the tool LanguageCrawl, which allows Natural Language Processing (NLP) researchers to easily build web-scale corpora using the Common Crawl Archive—an open repository of web crawl information, which contains petabytes of data. We present three use cases in the course of this work: filtering of Polish websites, the construction of n-gram corpora and the training of a continuous skipgram language model with hierarchical softmax. Each of them has been implemented within the LanguageCrawl toolkit, with the possibility to adjust specified language and n-gram ranks. This paper focuses particularly on high computing efficiency by applying highly concurrent multitasking. Our tool utilizes effective libraries and design. LanguageCrawl has been made publicly available to enrich the current set of NLP resources. We strongly believe that our work will facilitate further NLP research, especially in under-resourced languages, in which the lack of appropriately-sized corpora is a serious hindrance to applying data-intensive methods, such as deep neural networks.

Publisher

Springer Science and Business Media LLC

Subject

Library and Information Sciences,Linguistics and Language,Education,Language and Linguistics

Link

https://link.springer.com/content/pdf/10.1007/s10579-021-09551-7.pdf

Reference38 articles.

1. Banón, M., Chen, P., Haddow, B., Heafield, K., Hoang, H., Espla-Gomis, M., Forcada, M.L., Kamran, A., Kirefu, F., & Koehn, P., et al. (2020). Paracrawl: Web-scale acquisition of parallel corpora. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (pp. 4555–4567).

2. Bojanowski, P., Grave, E., Joulin, A., & Mikolov, T. (2017). Enriching word vectors with subword information. Transactions of the Association for Computational Linguistics, 5, 135–146.

3. Buck, C., Heafield, K., & van Ooyen, B. (2014). N-gram counts and language models from the common crawl. In Proceedings of the Language Resources and Evaluation Conference. Reykjavk, Icelandik, Iceland.

4. Crouse, S., Nagel, S., Elbaz, G., & Malamud, C. (2008). Common Crawl Foundation. http://commoncrawl.org

5. Ginter, F., & Kanerva, J. (2014). Fast training of word2vec representations using n-gram corpora. In: E. Volodina, L. Borin, I. Pilán (eds.) Linköping Electronic Conference Proceedings. Uppsala University.

Cited by 4 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Why do Words with Negative Connotations Still Exist? A Corpus-Based Analysis of the Words ‘Handicapped’, ‘Diffable’, and ‘Disability’;Rupkatha Journal on Interdisciplinary Studies in Humanities;2023-12-19

2. AI to Train AI: Using ChatGPT to Improve the Accuracy of a Therapeutic Dialogue System;Electronics;2023-11-18

3. Enhanced Emotion and Sentiment Recognition for Empathetic Dialogue System Using Big Data and Deep Learning Methods;Computational Science – ICCS 2023;2023

4. A Systematic Review on Machine Learning and Deep Learning Models for Electronic Information Security in Mobile Networks;Sensors;2022-03-04