While the stereo format is the most common representation for audio in general, it cannot provide a complete immersive audio experience due to its lack of surround sound. For this reason, multi-channel configurations such as 5.1 or 7.1 are used in cinema or home studios, offering an improved sonic sensation. However, converting a mono or stereophonic sound to multi-channel is often done manually by sound engineers and requires laborious work and effort. This paper presents an automatic algorithm for upmixing based on neural networks that can translate any stereo audio such as movies, music, or animated films, into a 5.1 or even 7.1 sound. Combining novel models for Source Separation and Primary-Ambient Extraction, our proposed method can reach near human-made upmixing results, evaluated subjectively using a MUSHRA-inspired listening test with 50 subjects. Furthermore, unlike other industrial approaches, our strategy is based on a mathematical formulation that ensures a near lossless round-trip conversion back to stereo, by downmixing the obtained multi-channel signal and comparing it with the original stereo, using SDR, SI-SDR, and MSE metrics.
