Utility-Theoretic Ranking for Semiautomated Text Classification-Reference-Cited by-同舟云学术

Utility-Theoretic Ranking for Semiautomated Text Classification

Published:2015-07-27 Issue:1 Volume:10 Page:1-32
ISSN:1556-4681
Container-title:ACM Transactions on Knowledge Discovery from Data
language:en
Short-container-title:ACM Trans. Knowl. Discov. Data

Author:

Berardi Giacomo¹,Esuli Andrea¹,Sebastiani Fabrizio²

Affiliation:

1. Italian National Council of Research

2. Qatar Computing Research Institute, Doha, Qatar

Abstract

Semiautomated Text Classification (SATC) may be defined as the task of ranking a set D of automatically labelled textual documents in such a way that, if a human annotator validates (i.e., inspects and corrects where appropriate) the documents in a top-ranked portion of D with the goal of increasing the overall labelling accuracy of D , the expected increase is maximized. An obvious SATC strategy is to rank D so that the documents that the classifier has labelled with the lowest confidence are top ranked. In this work, we show that this strategy is suboptimal. We develop new utility-theoretic ranking methods based on the notion of validation gain , defined as the improvement in classification effectiveness that would derive by validating a given automatically labelled document. We also propose a new effectiveness measure for SATC-oriented ranking methods, based on the expected reduction in classification error brought about by partially validating a list generated by a given ranking method. We report the results of experiments showing that, with respect to the baseline method mentioned earlier, and according to the proposed measure, our utility-theoretic ranking methods can achieve substantially higher expected reductions in classification error.

Publisher

Association for Computing Machinery (ACM)

Subject

General Computer Science

Link

https://dl.acm.org/doi/pdf/10.1145/2742548

Reference45 articles.

1. Incremental relevance feedback

2. A utility-theoretic ranking method for semi-automated text classification

3. Optimising human inspection work in automated verbatim coding

4. Dynamic ranked retrieval

Cited by 5 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Simulation of Big Data Order-Preserving Matching and Retrieval Model Based on Deep Learning;2023 International Conference on Power, Electrical Engineering, Electronics and Control (PEEEC);2023-09-25

2. Modifiers of Liver-Related Manifestation in the Course of NAFLD;Current Pharmaceutical Design;2020-04-24

3. Jointly Minimizing the Expected Costs of Review for Responsiveness and Privilege in E-Discovery;ACM Transactions on Information Systems;2019-01-31

4. Mining Domain Similarity to Enhance Digital Indexing;Proceedings of the 9th International Conference on Management of Digital EcoSystems;2017-11-07

5. From classification to quantification in tweet sentiment analysis;Social Network Analysis and Mining;2016-04-12