Audio-visual Sound Event Localization and Detection with Source Distance Estimation (AV 3D SELD) is designed to use synchronized audio-visual streams to simultaneously identify event categories, 3D spatial positions, and temporal intervals. In real-world scenarios, the visual modality provides rich spatiotemporal semantic cues regarding sound sources. Although recent AV 3D SELD studies have attempted to introduce visual depth cues, they still struggle to effectively exploit depth-aware visual geometry for distance estimation and cross-modal alignment. Furthermore, the absence of effective mechanisms to distill sound-related visual regions often introduces redundant background information, leading to performance degradation. To address these issues, we propose a novel framework called Spatial Semantic-Guided Network (SSGNet). Specifically, we introduce a Spatial Semantic Guidance Loss (SSGL), which guides the model to focus on sound-source-related spaces, effectively suppressing background interference. Additionally, we incorporate a spatial geometric prior to improve the association between sound sources and their visual locations, thereby boosting cross-modal fusion performance. Extensive experimental results on both DCASE 2023 and 2024 Challenge SELD tasks demonstrate that our method significantly outperforms existing SELD approaches.
