Data Sorting Influence on Short Text Manual Labeling Quality for Hierarchical Classification-Reference-Cited by-同舟云学术

Data Sorting Influence on Short Text Manual Labeling Quality for Hierarchical Classification

Published:2024-04-07 Issue:4 Volume:8 Page:41
ISSN:2504-2289
Container-title:Big Data and Cognitive Computing
language:en
Short-container-title:BDCC

Author:

Narushynska Olga¹^ORCID,Teslyuk Vasyl¹^ORCID,Doroshenko Anastasiya¹^ORCID,Arzubov Maksym¹^ORCID

Affiliation:

1. Department of Automated Control Systems, Lviv Polytechnic National University, 79013 Lviv, Ukraine

Abstract

The precise categorization of brief texts holds significant importance in various applications within the ever-changing realm of artificial intelligence (AI) and natural language processing (NLP). Short texts are everywhere in the digital world, from social media updates to customer reviews and feedback. Nevertheless, short texts’ limited length and context pose unique challenges for accurate classification. This research article delves into the influence of data sorting methods on the quality of manual labeling in hierarchical classification, with a particular focus on short texts. The study is set against the backdrop of the increasing reliance on manual labeling in AI and NLP, highlighting its significance in the accuracy of hierarchical text classification. Methodologically, the study integrates AI, notably zero-shot learning, with human annotation processes to examine the efficacy of various data-sorting strategies. The results demonstrate how different sorting approaches impact the accuracy and consistency of manual labeling, a critical aspect of creating high-quality datasets for NLP applications. The study’s findings reveal a significant time efficiency improvement in terms of labeling, where ordered manual labeling required 760 min per 1000 samples, compared to 800 min for traditional manual labeling, illustrating the practical benefits of optimized data sorting strategies. Comparatively, ordered manual labeling achieved the highest mean accuracy rates across all hierarchical levels, with figures reaching up to 99% for segments, 95% for families, 92% for classes, and 90% for bricks, underscoring the efficiency of structured data sorting. It offers valuable insights and practical guidelines for improving labeling quality in hierarchical classification tasks, thereby advancing the precision of text analysis in AI-driven research. This abstract encapsulates the article’s background, methods, results, and conclusions, providing a comprehensive yet succinct study overview.

Publisher

MDPI AG

Link

https://www.mdpi.com/2504-2289/8/4/41/pdf

Reference51 articles.

1. Automatic Content Analysis of Social Media Short Texts: Scoping Review of Methods and Tools;Costa;Computer Supported Qualitative Research,2020

2. Chat2VIS: Generating Data Visualizations via Natural Language Using ChatGPT, Codex and GPT-3 Large Language Models;Maddigan;IEEE Access,2023

3. Zhou, X., Wu, T., Chen, H., Yang, Q., and He, X. (2019, January 16–19). Automatic Annotation of Text Classification Data Set in Specific Field Using Named Entity Recognition. Proceedings of the 2019 IEEE 19th International Conference on Communication Technology (ICCT), Xi’an, China.

4. Doroshenko, A., and Tkachenko, R. (2018, January 11–14). Classification of Imbalanced Classes Using the Committee of Neural Networks. Proceedings of the 2018 IEEE 13th International Scientific and Technical Conference on Computer Sciences and Information Technologies (CSIT), Lviv, Ukraine.

5. Chang, C.-M., Mishra, S.D., and Igarashi, T. (2019, January 14–18). A Hierarchical Task Assignment for Manual Image Labeling. Proceedings of the 2019 IEEE Symposium on Visual Languages and Human-Centric Computing (VL/HCC), Memphis, TN, USA.

Cited by 1 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. An Efficient Algorithm for Sorting and Duplicate Elimination by Using Logarithmic Prime Numbers;Big Data and Cognitive Computing;2024-08-23