Text classification by CEFR levels using machine learning methods and BERT language model-Reference-Cited by-同舟云学术

Text classification by CEFR levels using machine learning methods and BERT language model

Published:2023-09-17 Issue:3 Volume:30 Page:202-213
ISSN:2313-5417
Container-title:Modeling and Analysis of Information Systems
language:
Short-container-title:Model. anal. inf. sist.

Author:

Lagutina Nadezhda S.¹^ORCID,Lagutina Ksenia V.¹^ORCID,Brederman Anastasya M.¹^ORCID,Kasatkina Natalia N.¹^ORCID

Affiliation:

1. P.G. Demidov Yaroslavl State University

Abstract

This paper presents a study of the problem of automatic classification of short coherent texts (essays) in English according to the levels of the international CEFR scale. Determining the level of text in natural language is an important component of assessing students knowledge, including checking open tasks in e-learning systems. To solve this problem, vector text models were considered based on stylometric numerical features of the character, word, sentence structure levels. The classification of the obtained vectors was carried out by standard machine learning classifiers. The article presents the results of the three most successful ones: Support Vector Classifier, Stochastic Gradient Descent Classifier, LogisticRegression. Precision, recall and F-score served as quality measures. Two open text corpora, CEFR Levelled English Texts and BEA-2019, were chosen for the experiments. The best classification results for six CEFR levels and sublevels from A1 to C2 were shown by the Support Vector Classifier with F-score 67 % for the CEFR Levelled English Texts. This approach was compared with the application of the BERT language model (six different variants). The best model, bert-base-cased, provided the F-score value of 69 %. The analysis of classification errors showed that most of them are between neighboring levels, which is quite understandable from the point of view of the domain. In addition, the quality of classification strongly depended on the text corpus, that demonstrated a significant difference in F-scores during application of the same text models for different corpora. In general, the obtained results showed the effectiveness of automatic text level detection and the possibility of its practical application.

Publisher

P.G. Demidov Yaroslavl State University

Subject

General Medicine

Reference22 articles.

1. E. del Gobbo, A. Guarino, B. Cafarelli, L. Grilli, and P. Limone, “Automatic evaluation of open-ended questions for online learning. A systematic mapping,” Studies in Educational Evaluation, vol. 77, p. 101258, 2023.

2. N. V. Galichev and P. S. Shirogorodskaya, “Problema avtomaticheskogo izmereniya slozhnyh konstruktov cherez otkrytye zadaniya,” in HXI Mezhdunarodnaya nauchno-prakticheskaya konferenciya molodyh issledovatelej obrazovaniya, 2022, pp. 695–697.

3. L. E. Adamova, O. V. Surikova, I. G. Bulatova, and O. O. Varlamov, “Application of the mivar expert system to evaluate the complexity of texts,” News of the Kabardin-Balkar scientific center of RAS, no. 2, pp. 11–29, 2021.

4. D. Ramesh and S. K. Sanampudi, “An automated essay scoring systems: a systematic literature review,” Artificial Intelligence Review, vol. 55, no. 3, pp. 2495–2527, 2022.

5. K. P. Yancey, G. Laflair, A. Verardi, and J. Burstein, “Rating Short L2 Essays on the CEFR Scale with GPT-4,” in Proceedings of the 18th Workshop on Innovative Use of NLP for Building Educational Applications (BEA 2023), 2023, pp. 576–584.