DeepInspect: A Black-box Trojan Detection and Mitigation Framework for Deep Neural Networks-Reference-Cited by-同舟云学术

DeepInspect: A Black-box Trojan Detection and Mitigation Framework for Deep Neural Networks

Published:2019-08 Issue: Volume: Page:
ISSN:
Container-title:Proceedings of the Twenty-Eighth International Joint Conference on Artificial Intelligence
language:
Short-container-title:

Author:

Chen Huili¹,Fu Cheng¹,Zhao Jishen¹,Koushanfar Farinaz¹

Affiliation:

1. University of California, San Diego

Abstract

Deep Neural Networks (DNNs) are vulnerable to Neural Trojan (NT) attacks where the adversary injects malicious behaviors during DNN training. This type of ‘backdoor’ attack is activated when the input is stamped with the trigger pattern specified by the attacker, resulting in an incorrect prediction of the model. Due to the wide application of DNNs in various critical fields, it is indispensable to inspect whether the pre-trained DNN has been trojaned before employing a model. Our goal in this paper is to address the security concern on unknown DNN to NT attacks and ensure safe model deployment. We propose DeepInspect, the first black-box Trojan detection solution with minimal prior knowledge of the model. DeepInspect learns the probability distribution of potential triggers from the queried model using a conditional generative model, thus retrieves the footprint of backdoor insertion. In addition to NT detection, we show that DeepInspect’s trigger generator enables effective Trojan mitigation by model patching. We corroborate the effectiveness, efficiency, and scalability of DeepInspect against the state-of-the-art NT attacks across various benchmarks. Extensive experiments show that DeepInspect offers superior detection performance and lower runtime overhead than the prior work.

Publisher

International Joint Conferences on Artificial Intelligence Organization

Cited by 117 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Towards robustness evaluation of backdoor defense on quantized deep learning models;Expert Systems with Applications;2024-12

2. OCGEC: One-class Graph Embedding Classification for DNN Backdoor Detection;2024 International Joint Conference on Neural Networks (IJCNN);2024-06-30

3. Robust and privacy-preserving collaborative training: a comprehensive survey;Artificial Intelligence Review;2024-06-20

4. Nightshade: Prompt-Specific Poisoning Attacks on Text-to-Image Generative Models;2024 IEEE Symposium on Security and Privacy (SP);2024-05-19

5. TrojanPuzzle: Covertly Poisoning Code-Suggestion Models;2024 IEEE Symposium on Security and Privacy (SP);2024-05-19