Skip to main content

Toward Harmless Ownership Verification for Automatic Speech Recognition via Synthetic Lexical Domains

By
Hanbo Cai; Pengcheng Zhang; Yan Xiao; Shunhui Ji; Mingxuan Xiao

Automatic speech recognition (ASR) has been widely applied in intelligent assistants and voice services, and high-performance ASR models trained on large-scale speech data have become important commercial assets. However, the protection and verification of ASR model ownership face significant challenges. Most existing watermarking methods based on backdoor mechanisms rely on controllable misclassification, i.e., trigger-induced erroneous outputs for ownership verification, which may compromise model reliability under normal operating conditions and create abnormal input-output mappings, thereby introducing additional security vulnerabilities. Therefore, instead of inducing erroneous predictions on in-domain inputs, we propose a harmless ASR model ownership verification method that embeds watermarks in an external and non-interfering recognition behavior. Specifically, we use a large language model to generate new words, and combine them with text-to-speech technology to create an external synthetic lexical domain. This domain generates speech samples with minimal overlap with the original domain, which serve as the watermark. The model is then fine-tuned to learn these samples. In the verification phase, we also design a hypothesis testing method to perform ownership verification on the synthetic samples input by the model owner. Extensive experiments show that our method achieves stable and efficient ownership verification across various mainstream ASR models and datasets, with an average watermark success rate of over 95%, without degrading the original task performance. It is also effective in resisting fine-tuning, pruning, adaptive filtering and distortion attacks.

Read on IEEE Xplore