Test Selection for Deep Learning Systems-Reference-Cited by-同舟云学术

Test Selection for Deep Learning Systems

Published:2021-04-30 Issue:2 Volume:30 Page:1-22
ISSN:1049-331X
Container-title:ACM Transactions on Software Engineering and Methodology
language:en
Short-container-title:ACM Trans. Softw. Eng. Methodol.

Author:

Ma Wei¹,Papadakis Mike¹,Tsakmalis Anestis¹^ORCID,Cordy Maxime¹,Traon Yves Le¹

Affiliation:

1. University of Luxembourg, Luxembourg

Abstract

Testing of deep learning models is challenging due to the excessive number and complexity of the computations involved. As a result, test data selection is performed manually and in an ad hoc way. This raises the question of how we can automatically select candidate data to test deep learning models. Recent research has focused on defining metrics to measure the thoroughness of a test suite and to rely on such metrics to guide the generation of new tests. However, the problem of selecting/prioritising test inputs (e.g., to be labelled manually by humans) remains open. In this article, we perform an in-depth empirical comparison of a set of test selection metrics based on the notion of model uncertainty (model confidence on specific inputs). Intuitively, the more uncertain we are about a candidate sample, the more likely it is that this sample triggers a misclassification. Similarly, we hypothesise that the samples for which we are the most uncertain are the most informative and should be used in priority to improve the model by retraining. We evaluate these metrics on five models and three widely used image classification problems involving real and artificial (adversarial) data produced by five generation algorithms. We show that uncertainty-based metrics have a strong ability to identify misclassified inputs, being three times stronger than surprise adequacy and outperforming coverage-related metrics. We also show that these metrics lead to faster improvement in classification accuracy during retraining: up to two times faster than random selection and other state-of-the-art metrics on all models we considered.

Funder

FNR

Publisher

Association for Computing Machinery (ACM)

Subject

Software

Link

https://dl.acm.org/doi/pdf/10.1145/3417330

Reference40 articles.

1. Is mutation an appropriate tool for testing experiments?

2. Adversarial Examples Are Not Easily Detected

3. Towards Evaluating the Robustness of Neural Networks

Cited by 65 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Datactive: Data Fault Localization for Object Detection Systems;Proceedings of the 33rd ACM SIGSOFT International Symposium on Software Testing and Analysis;2024-09-11

2. Test Selection for Deep Neural Networks using Meta-Models with Uncertainty Metrics;Proceedings of the 33rd ACM SIGSOFT International Symposium on Software Testing and Analysis;2024-09-11

3. Unveiling Code Pre-Trained Models: Investigating Syntax and Semantics Capacities;ACM Transactions on Software Engineering and Methodology;2024-08-26

4. DeepSense: test prioritization for neural network based on multiple mutation and manifold spatial distribution;Evolutionary Intelligence;2024-07-30

5. Neuron importance-aware coverage analysis for deep neural network testing;Empirical Software Engineering;2024-07-25