Existing neural-network-based multichannel speech enhancement methods often rely on magnitude or real-imaginary modeling, but pay insufficient attention to explicit phase reconstruction and the effective exploitation of spatial cues derived from inter-channel phase difference (IPD) features, which limits overall enhancement performance. To address these issues, this paper proposes a novel Multichannel Speech Enhancement model with Spatial-Aware explicit Phase Estimation (SAPE-MSE). The SAPE-MSE adopts a magnitude-phase encoder-decoder architecture. The encoder encodes the multichannel degraded magnitude and phase spectra, followed by deep processing with improved time-frequency Transformers. Parallel decoders then estimate the clean magnitude and phase spectra at the reference channel, followed by inverse transformation to reconstruct the clean speech. Importantly, the phase decoding branch is guided by IPD-derived spatial cues and explicitly estimates the clean phase spectrum via circular array spatial processing and parallel phase estimation. Experimental results on a synthetic eight-channel dataset show that SAPE-MSE achieves competitive performance, obtaining the best PESQ of 3.58 as well as the best STOI, ESTOI, and DNSMOS among the evaluated methods.
