Simple or Complex? Learning to Predict Readability of Bengali Texts-Reference-Cited by-同舟云学术

Simple or Complex? Learning to Predict Readability of Bengali Texts

Published:2021-05-18 Issue:14 Volume:35 Page:12621-12629
ISSN:2374-3468
Container-title:Proceedings of the AAAI Conference on Artificial Intelligence
language:
Short-container-title:AAAI

Author:

Chakraborty Susmoy,Nayeem Mir Tafseer,Ahmad Wasi Uddin

Abstract

Determining the readability of a text is the first step to its simplification. In this paper, we present a readability analysis tool capable of analyzing text written in the Bengali language to provide in-depth information on its readability and complexity. Despite being the 7th most spoken language in the world with 230 million native speakers, Bengali suffers from a lack of fundamental resources for natural language processing. Readability related research of the Bengali language so far can be considered to be narrow and sometimes faulty due to the lack of resources. Therefore, we correctly adopt document-level readability formulas traditionally used for U.S. based education system to the Bengali language with a proper age-to-age comparison. Due to the unavailability of large-scale human-annotated corpora, we further divide the document-level task into sentence-level and experiment with neural architectures, which will serve as a baseline for the future works of Bengali readability prediction. During the process, we present several human-annotated corpora and dictionaries such as a document-level dataset comprising 618 documents with 12 different grade levels, a large-scale sentence-level dataset comprising more than 96K sentences with simple and complex labels, a consonant conjunct count algorithm and a corpus of 341 words to validate the effectiveness of the algorithm, a list of 3,396 easy words, and an updated pronunciation dictionary with more than 67K words. These resources can be useful for several other tasks of this low-resource language.

Publisher

Association for the Advancement of Artificial Intelligence (AAAI)

Subject

General Medicine

Cited by 8 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Analisis Keterbacaan Teks Buku Ajar Bahasa Indonesia SMP Kelas 9 Menggunakan Formula Grafik Fry;Pubmedia Jurnal Penelitian Tindakan Kelas Indonesia;2024-05-17

2. Multisensory computer-based system for teaching sentence reading in Hindi and Bangla to children with dyslexia;Technology and Disability;2023-12-27

3. A Machine Learning-Based Readability Model for Gujarati Texts;ACM Transactions on Asian and Low-Resource Language Information Processing;2023-12-21

4. Navigating Bengali Linguistics: Insights from Machine and Deep Learning Perspectives for Categorization of Sentences;2023 26th International Conference on Computer and Information Technology (ICCIT);2023-12-13

5. Beyond Words: Unraveling Text Complexity with Novel Dataset and A Classifier Application;2023 26th International Conference on Computer and Information Technology (ICCIT);2023-12-13