Skip to main content

EquilSE: Time-Invariant Generative Speech Enhancement via Equilibrium Matching

By
Hao Liang; Wei Liu; Gongping Huang; Shoji Makino

Diffusion- and flow-based generative models have achieved strong performance in speech enhancement, but they typically rely on an explicit time variable to specify the generation stage. Since the noisy speech condition provides a strong reference to the underlying clean speech, enhancement progress could be reflected by the relation between the current sample and the noisy speech. This raises the question of whether the enhancement stage needs to be specified to the network through an explicit time variable. To study this question, we introduce EquilSE, an equilibrium matching-based framework that models generative speech enhancement with a time-invariant conditional vector field and formulates inference as an optimization process. By removing explicit time input and introducing an equilibrium-inducing target, EquilSE achieves comparable performance to time-dependent baselines under few-step inference and provides improvements in the single-step regime. Furthermore, EquilSE remains competitive under cross-domain evaluation.

Read on IEEE Xplore