Uncertainty-based active learning with instability estimation for text classification-Reference-Cited by-同舟云学术

Uncertainty-based active learning with instability estimation for text classification

Published:2012-02 Issue:4 Volume:8 Page:1-21
ISSN:1550-4875
Container-title:ACM Transactions on Speech and Language Processing
language:en
Short-container-title:ACM Trans. Speech Lang. Process.

Author:

Zhu Jingbo¹,Ma Matthew²

Affiliation:

1. Northeastern University, China

2. Scientific Works, Princeton, NJ

Abstract

This article deals with pool-based active learning with uncertainty sampling. While existing uncertainty sampling methods emphasize selection of instances near the decision boundary to increase the likelihood of selecting informative examples, our position is that this heuristic is a surrogate for selecting examples for which the current learning algorithm iteration is likely to misclassify. To more directly model this intuition, this article augments such uncertainty sampling methods and proposes a simple instability -based selective sampling approach to improving uncertainty-based active learning, in which the instability degree of each unlabeled example is estimated during the learning process. Experiments on seven evaluation datasets show that instability-based sampling methods can achieve significant improvements over the traditional uncertainty sampling method. In terms of the average percentage of actively selected examples required for the learner to achieve 99% of its performance when training on the entire dataset, instability sampling and sampling by instability and density methods achieve better effectiveness in annotation cost reduction than random sampling and traditional entropy-based uncertainty sampling. Our experimental results have also shown that instability-based methods yield no significant improvement for active learning with SVMs when a popular sigmoidal function is used to transform SVM outputs to posterior probabilities.

Funder

Research Grants Council, University Grants Committee, Hong Kong

National Natural Science Foundation of China

Publisher

Association for Computing Machinery (ACM)

Subject

Computational Mathematics,Computer Science (miscellaneous)

Link

https://dl.acm.org/doi/pdf/10.1145/2093153.2093154

Reference46 articles.

1. Online choice of active learning algorithms;Yoram B.;J. Mach. Learn. Res.,2004

Cited by 7 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Understanding Human-side Impact of Sampling Image Batches in Subjective Attribute Labeling;Proceedings of the ACM on Human-Computer Interaction;2021-10-13

2. Looking Back on the Past: Active Learning with Historical Evaluation Results;IEEE Transactions on Knowledge and Data Engineering;2021

3. LeSSA: A Unified Framework based on Lexicons and Semi-Supervised Learning Approaches for Textual Sentiment Classification;Applied Sciences;2019-12-17

4. Modeling of learning curves with applications to POS tagging;Computer Speech & Language;2017-01

5. Combination of active learning and self-training for cross-lingual sentiment classification with density analysis of unlabelled samples;Information Sciences;2015-10