Skip to main content

Unveiling the Backdoor’s Suppression Effect for Backdoor Defence

By
Hong Zhu; Kai Chen; Shengzhi Zhang

Deep neural networks (DNNs) are vulnerable to backdoor/Trojan attacks, posing a severe threat to security-critical applications. In this paper, we investigate a scenario in which two different backdoors are successively embedded into one model and report that the later-embedded backdoor has a “suppression” effect on the earlier-embedded backdoor. On the basis of this suppression effect, we intentionally embed a special backdoor, named Safedoor, into suspect DNNs that may have already been infected by backdoors. In this way, our later-embedded Safedoor defends against earlier-embedded backdoor attacks. We also propose an efficient Safedoor embedding algorithm that makes the defence feasible with only a small amount of data. We evaluate Safedoor using eight representative backdoor attack approaches, and the results demonstrate its efficacy in defeating various backdoor attacks. Specifically, Safedoor reduces the attack success rate from 97.3% to 2.2% on average, outperforming state-of-the-art methods by 39.8%. It also maintains high accuracy on clean data, outperforming state-of-the-art methods by 1.8%. Furthermore, Safedoor introduces only 0.32% overhead to the model inference time, which is negligible in most cases.

Read on IEEE Xplore