Skip to main content

Cyclostationarity Analysis as a Complement to Self-Supervised Representations for Speech Deepfake Detection

By
Cemal Hanilçi; Md. Sahidullah; Tomi H. Kinnunen

Speech deepfake detection (SDD) is essential for maintaining trust in voice-driven technologies and digital media. Although recent SDD systems increasingly rely on self-supervised learning (SSL) representations that capture rich contextual information, complementary signal-driven acoustic features remain important for modeling fine-grained structural properties of speech.Most existing acoustic front ends are based on time-frequency representations, which do not fully exploit higher-order spectral dependencies inherent in speech signals. We introduce a cyclostationarity-inspired acoustic feature extraction framework for SDD based on spectral correlation density (SCD). The proposed features model periodic statistical structures in speech by capturing spectral correlations between frequency components. In particular, we investigate whether cyclostationary representations provide complementary information beyond conventional time-frequency features and modern SSL embeddings. To this end, we introduce temporally structured SCD representations that preserve the evolution of spectral and cyclic-frequency correlations over time. Their effectiveness is evaluated using multiple SDD architectures, including conventional neural networks, SSL-based embedding systems, and hybrid fusion models. Experiments on ASVspoof 2019 LA, ASVspoof 2021 DF, and ASVspoof 5 demonstrate that higher-order spectral correlations provide complementary information to SSL embeddings in several challenging conditions. In particular, fusion of SSL and SCD embeddings reduces the equal error rate on ASVspoof 2019 LA from 8.28% to 0.98%, and from 15.81% to 14.80% on the challenging ASVspoof 5 dataset. These findings establish higher-order cyclostationary spectral correlations as a complementary source of information for modern SDD, while providing new signal-processing insight into the statistical differences between natural and synthetic speech.

Read on IEEE Xplore