Skip to main content

Stealthy Physical Adversarial Attacks on Speaker Recognition via Near-Ultrasonic Perturbations

By
Man Zhou; Junhui Yang; Xinyuan Chen; Zhehan Chen; Xiangsheng Zeng; Jun Feng; Zhengxiong Li; Qi Li; Qian Wang

Speaker recognition systems (SRS) based on deep neural networks (DNNs) are widely deployed in security-critical applications. However, they remain vulnerable to attacks that manipulate audio inputs to spoof user identities. Existing physical attacks against SRS face critical trade-offs: audible-band approaches compromise stealthiness or robustness, while ultrasound-based methods achieve inaudibility but require specialized hardware. In this paper, we propose AdvNup, a stealthy physical adversarial attack that leverages near-ultrasonic perturbations to bypass human hearing by using commercial off-the-shelf (COTS) speakers. To overcome limited bandwidth, AdvNup leverages single-sideband modulation with low-pass filtering, enabling adversarial optimization within the constrained bandwidth. Then we introduce the nonlinear frequency response (NFR) through a step signal demodulation analysis to mathematically characterize the nonlinear demodulation, ensuring adversarial integrity during physical transmission. Furthermore, we improve robustness by extending NFR modeling to spatial variations and implementing time-frequency masking to mitigate environmental interference. Comprehensive simulated and physical experiments demonstrate the robustness and effectiveness of AdvNup. It can deceive state-of-the-art SR models with a 100% attack success rate (ASR) for closed-set identification (CSI) and more than 90% ASR for open-set identification (OSI) in a targeted manner, significantly outperforming existing stealthy physical adversarial attacks.

Read on IEEE Xplore