Self-supervised audio models are typically pre-trained on large collections of real-world recordings, and the resulting representations may be shaped by the acoustic coverage, recording conditions, and dataset-specific correlations of the source corpus. We investigate whether procedurally generated audio can provide a complementary acoustic prior for masked audio pre-training. To this end, we propose PRP, a procedural-to-real pre-training framework that combines a multi-source procedural audio synthesizer with a Transformer-based masked autoencoder. The synthesizer generates multi-event waveforms containing diverse harmonic, resonant, frequency-modulated, transient, and stochastic structures, together with randomized temporal envelopes, spectral coloration, and channel perturbations. PRP parameterizes the fraction of pre-training updates drawn from procedural audio, enabling matched comparisons among procedural-only, real-only, and mixed source plans under the same architecture, reconstruction objective, sample budget, and number of optimization steps. Using an equal procedural–real source ratio, PRP-Mix achieves 92.00% accuracy on ESC-50, 89.84% on UrbanSound8K, 91.50% on GTZAN, 97.37% on Speech Commands v2, and 0.5979 mAP on FSD50 K, outperforming the matched real-only masked autoencoder on all five benchmarks. PRP-Mix also provides stronger average transfer in limited-label experiments on ESC-50 and FSD50 K and outperforms a matched 50:50 control based on a simpler synthetic generator. Analysis of frozen representations further shows that frequency, relative intensity, and onset position remain linearly decodable, although these controls are encoded in a distributed manner and the analysis does not establish causal disentanglement. Overall, the results demonstrate that procedurally generated audio can complement real-domain masked pre-training under a fixed optimization budget, while the preferred procedural–real ratio remains dependent on the downstream task.
