Voice Activity Detection (VAD) is a fundamental component that supports a wide range of audio/speech processing applications. Numerous studies have addressed VAD using a single microphone or a compact microphone array, yet their performance remains limited in distant scenarios. In this paper, we develop a novel data-driven VAD framework based on Wireless Acoustic Sensor Networks (WASNs), which are expected to serve as the next-generation platform for audio/speech processing. The network architecture consists of four key components: Local Representation Learning (LRL), Quantization Encoding (QE), Channel Aggregation (CA), and Temporal Modeling (TM). Among them, LRL and QE are deployed at local nodes. Specifically, LRL extracts latent features from recordings, enabling nodes to offload part of the computations from the Fusion Center (FC). Then, QE dynamically encodes latent features into a few discrete tokens, which can reduce a large amount of data transmission between nodes and the FC. CA and TM are placed at the FC, where the former fuses multichannel data received from nodes and the latter captures inter-frame dependencies and produces VAD results. Furthermore, we design a two-stage training strategy to enhance stability and convergence during training. Numerical experiments demonstrate that the proposed method outperforms existing model-driven or data-driven VAD approaches over WASNs.
