Zero-shot voice conversion (ZSVC) enables timbre transfer from unseen speakers without fine-tuning, raising serious risks of audio misuse and copyright infringement. To protect potential victims of such misuse, we propose VocaLock, a watermark-based system for detecting forged audio and attributing speaker timbre under ZSVC attacks. VocaLock embeds a single watermark into the user’s audio and employs two decoders trained with distinct robustness objectives: a robust decoder for timbre attribution and a semi-robust decoder for fake audio detection. When ZSVC is used to attribute a manipulated content to a victim, the robust decoder successfully extracts the watermark, while the semi-robust decoder fails. To achieve watermark robustness against the structural distortions introduced by VC, we embed the watermark in the short-time Fourier transform (STFT) spectrum and introduce a post-VC loss that aligns watermark embedding with the timbre feature space, allowing watermark traces to transfer alongside timbre features into forged audio. Robustness and generalization to black-box VC models are further enhanced through training under cross-domain distortions. Additional architectural and training strategies are employed to balance robustness and audio fidelity. Experimental results demonstrate that VocaLock exhibits strong robustness and generalization against various ZSVC models and common post-processing operations, while maintaining good audio fidelity.
