Towards building a Urdu Language Corpus using Common Crawl-Reference-Cited by-同舟云学术

Towards building a Urdu Language Corpus using Common Crawl

Published:2020-08-31 Issue:2 Volume:39 Page:2445-2455
ISSN:1064-1246
Container-title:Journal of Intelligent & Fuzzy Systems
language:
Short-container-title:IFS

Author:

Shafiq Hafiz Muhammad¹,Tahir Bilal¹,Mehmood Muhammad Amir¹

Affiliation:

1. Al-Khawarizmi Institute of Computer Science, University of Engineering and Technology, Lahore, Pakistan

Abstract

Urdu is the most popular language in Pakistan which is spoken by millions of people across the globe. While English is considered the dominant web content language, characteristics of Urdu language web content are still unknown. In this paper, we study the World-Wide-Web (WWW) by focusing on the content present in the Perso-Arabic script. Leveraging from the Common Crawl Corpus, which is the largest publicly available web content of 2.87 billion documents for the period of December 2016, we examine different aspects of Urdu web content. We use the Compact Language Detector (CLD2) for language detection. We find that the global WWW population has a share of 0.04% for Urdu web content with respect to document frequency. 70.9% of the top-level Urdu domains consist of . com, . org, and . info. Besides, urdulughat is the most dominating second-level domain. 40% of the domains are hosted in the United States while only 0.33% are hosted within Pakistan. Moreover, 25.68% web-pages have Urdu as primary language and only 11.78% of web-pages are exclusively in Urdu. Our Urdu corpus consists of 1.25 billion total and 18.14 million unique tokens. Furthermore, the corpus follows the Zipf’s law distribution. This Urdu Corpus can be used for text summarization, text classification, and cross-lingual information retrieval.

Publisher

IOS Press

Subject

Artificial Intelligence,General Engineering,Statistics and Probability

Reference1 articles.

1. Large-scale analysis of zipf's law in english texts;Moreno-Sanchez;Journal of PloS one,2016

Cited by 5 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Intent Detection in Urdu Queries using Fine-tuned BERT models;2022 16th International Conference on Open Source Systems and Technologies (ICOSST);2022-12-14

2. UBERT22: Unsupervised Pre-training of BERT for Low Resource Urdu Language;2022 16th International Conference on Open Source Systems and Technologies (ICOSST);2022-12-14

3. Improving User Intent Detection in Urdu Web Queries with Capsule Net Architectures;Applied Sciences;2022-11-21

4. Anbar: Collection and analysis of a large scale Urdu language Twitter corpus;Journal of Intelligent & Fuzzy Systems;2021-12-24

5. Corpulyzer: A Novel Framework for Building Low Resource Language Corpora;IEEE Access;2021