The aim of a backdoor attack is to manipulate the behavior of a DNN under trigger-activated conditions by poisoning the dataset. Most backdoor detectors can be regarded as feature clustering methods and struggle to detect trigger-injected samples owing to two factors: 1) Most detectors rely heavily on sufficient clean samples and deficient artificial priors to analyze the specific clustering, which leads to the necessity of rebuilding the analysis models when new attacks are detected; and 2) Modern backdoor features are usually coupled with benign features, which critically hinders the performance of unlabeled feature clustering detection methods. In this study, instead of using an unlabeled feature clustering framework, an Advanced Cross-attack Backdoor Detector (ACBD) is trained by solving a labeled binary classification task that is based on the commonality identified from the objective of the attacks, i.e., backdoor attacks lead victim models to classify the triggers disturbed by images into the target label (which is defined as the disturbance immunity of triggers). Specifically, our ACBD needs only a single class of the poisoned dataset (e.g., 1/100 of CIFAR-100) with two classic attacks to construct a small labeled dataset for training a small LSTM with 53 K parameters. We perturb the new dataset by a few clean images $(\leq 10)$ to drive our ACBD to learn the disturbance immunity. Our ACBD exhibits state-of-the-art (SOTA) detection performance with empirical cross-attack generalization on various hard-to-detect attacks, even with target labels that are different from the labels used in training. Given corresponding victim models, extensive test-time experiments reveal that our ACBD also achieves satisfactory detection performance on a test set with unseen triggers.
