Skip to main content

Generative Textual Adversarial Attack Through Extensible Compositional Perturbation via Reinforcement Learning for Policy Optimization

By
Hao-Tian Wu; Jiachen Xie; Jiankun Hu; Zhihong Tian

As textual adversarial attacks pose risks to natural language processing models, adversarial training is often utilized to improve robustness of these systems. In a strict black-box scenario, most query-based textual adversarial attack methods rely on fixed or limited perturbation policies and search in the perturbation space, making it difficult to balance attack success rate (ASR), query cost and textual quality of generated adversarial examples. To address this, we propose GECOMP, a generative textual adversarial attack method that relies on policy optimization by constructing an extensible library of compositional perturbations. By adopting a large language model (LLM) based generator to rewrite an input sequence, the outputs are perturbed according to the library under constraints of semantic similarity and edit magnitude. By employing a budget-aware reinforcement learning strategy for training with respect to feedback of a victim model, the LLM-based generator learns the most effective perturbation policy for adversarial example generation. Experimental results and performance comparisons on four public datasets and across five victim models show that, compared with ten baseline methods, GECOMP generates adversarial examples with higher ASRs and better quality by using fewer queries and less edit magnitude.

Read on IEEE Xplore