Skip to main content

Efficient Front-End Speech Enhancement for Robust ASR via Parallel Time–Band Mixing and Learned Observation Fusion

By
Xingyu Shen; Runze Wang; Wei-Ping Zhu; Benoit Champagne

Although front-end speech enhancement can improve automatic speech recognition (ASR) robustness, its practical deployment under frozen recognizers is often limited by within-block sequential dependencies (recurrent unrolling) and enhancement artifacts that amplify downstream errors. We propose an efficient band-split enhancement front end based on a Parallel Time–Band Mixer (PTBM), designed to capture intra-band temporal context and cross-band structure while facilitating parallel computation. To suppress artifacts that are harmful to ASR without development-set coefficient tuning, we introduce ASR-oriented learned observation fusion (LOF), which predicts frame–band fusion weights on the band-split representation and performs structured complex-spectrum fusion on the short-time Fourier transform (STFT) grid by expanding these weights to the corresponding frequency bins. During training, we optionally incorporate a lightweight auxiliary consistency objective based on connectionist temporal classification (CTC) from a frozen teacher, while keeping the downstream recognizers frozen. We evaluate the proposed method on a diverse suite of noisy and reverberant benchmark datasets and real recordings, including English and Mandarin settings, and further conduct analyses stratified by signal-to-noise ratio (SNR) and out-of-domain stress tests. Our front end consistently reduces WER/CER relative to strong enhancement and fusion-based baselines, while lowering front-end computation from 1.84 to 0.60 GMAC/s relative to a representative recurrent band-split baseline.

Read on IEEE Xplore