Most deep learning–based sound source localization methods rely on specific microphone array geometries, which require retraining for different configurations and lead to high cost. IPDnet is our previous work which addresses this limitation through pair-wise processing with a mean pooling scheme. However, processing narrow-band independently leads to high computational complexity. In this work, we propose IPDnet2A, an efficient model which improves localization performance while reducing computational complexity. IPDnet2A adopts oSpatialNet as the backbone to extract spatial features, and a frequency pooling mechanism is used to compress the frequency dimension, thus reducing computational cost. Additionally, an interaction module is designed to process the mean pair-wise representation, which further improves interaction between microphone pairs. Experiments on multiple datasets demonstrate that IPDnet2A achieves state-of-the-art localization performance while significantly reducing computational cost compared to IPDnet.
