Testing Ambisonics decoders
Over the past few weeks, I have systematically tested preferred methods for transposing1 recordings and works from Ambisonics to Dolby Atmos. The first step, and the subject of this post, is how to decode the Ambisonics sound field.
Still from the video.
I have used one of the scenes from Mülheim an der Ruhr, August 2013 for testing. The sound for this work was recorded with a SoundField SPS200 first-order Ambisonics microphone, and when revisiting the material, I start from an earlier encoded FuMa B-format version of the file, already synchronised to the video. It will be decoded to either 7.0.4 or 9.0.6, meaning that the four channels of the original recording are to provide eleven or fifteen channels of audio. As such, the original recording is under-specified in terms of spatial resolution for the number of channels required in the final outcome. From experience, I know that simply decoding in first order will lead to an unstable result, with phasing issues and more. But in recent years, several solutions have emerged for spatial upscaling to higher-order Ambisonics. Now was the time to test and compare a number of them.
I do not believe there is one and only one best way to do this. Different approaches will sound different, but that does not necessarily translate into “better or worse”. The differences open a creative and interpretative space, and, ultimately, the preferred approach will depend as much on artistic intent as on technological capability and affordance. An important question that emerges in the process is therefore “What are the quality criteria this time, with respect to the artistic project and the sound material I am dealing with here?”
Mülheim an der Ruhr, August 2013 is a series of audio-visual field recordings from suburban environments. I want the resulting Dolby Atmos transpositions1 to preserve a sense of place, giving the audience an impression of “being there”, immersed in sound environments from places that may initially seem bland, banal, and non-interesting, yet which ultimately reveal themselves as sonically and spatially rich, varied, and worth engaging with and caring for. This led me to define a more specific set of quality criteria that I search for when listening to the outcomes of the various approaches:
- A sense of place
- Continuity and fullness in the resulting sound field
- Spatial differentiation and articulation
- Sound field stability, avoiding phasing
- Avoiding overly strong source separation (avoiding that the sound sources within the sound scene jump from speaker to speaker with movement, and hence point out the specific locations of each speaker rather than give the illusion of a continuous sound field)
- Avoiding audible processing artefacts (digital, glitch, noise, spectral)
- Continuity and naturalness in the sound when soloing individual channels (avoiding single-channel artefacts, even if the overall reproduced sound might still sound convincing)
- Maintain spatial and sonic quality and integrity as far as possible when folding down from 9.0.6 or 7.0.4 to 7.1, 5.1, stereo and binaural
- CPU demand, for pragmatic reasons
Screenshot of the Reaper project with effect processing for decoding the first-order Ambisonics signal. Separate FX containers are used for each approach.
The screenshot above shows the Reaper project used for testing. All decoding approaches are configured on the same channel, each in a separate FX container, making it easy to switch between them and adjust the number of channels as needed for each decoder. Gain matching was applied to ensure consistent integrated levels regardless of which decoder was used.
Harpex
Harpex decoding.
Harpex has a preset offering direct decoding of the first-order signal to 7.0.4. However, 9.0.6 is not an option. When decoding to a horizontal-only surround, one can adjust the azimuth for each speaker and vary the emulated distance between the virtual microphones, and hence the amount of decorrelation between the speaker channels. These parameters are not available within the 7.0.4 preset. It should also be noted that while the 7.0 preset locates the rear speakers at ±135°, they seem to be located at ±150° in the 7.0.4 preset.
Overall, the resulting sound field feels convincing, but some artefacts are noticeable, and when soloing individual channels, they get pronounced. When first released, Harpex was ground-breaking, but as of 2026, other and better options are available when needing to decode to larger sets of speakers.
SPARTA plugins
The SPARTA set of plugins from Aalto University offers several approaches.
Sparta Compass decoding.
The compass_decoder does parametrically enhanced decoding up to 3rd order. Decoding straight from first-order Ambisonics did not sound convincing, with pronounced artefacts in individual channels. Alternatively, I upscaled to third order using ab Image Upscaler before decoding. This works better, but the sound field is perceived as unstable, with sources jumping between speakers depending on their direction of arrival, while still giving artefacts in individual channels. This decoder also seems CPU-heavy, causing playback glitches. The Diffuse to Direct and Linear to Parametric parameters can probably be fine-tuned for better results, but the conclusion is that this is not the preferred solution.
SPARTA HO-DirAC decoding.
The hodirac_decoder uses an alternative Higher-order Directional Audio Coding method for up to third-order input and also offers parameters for Diffuse to Direct and Linear to Parametric, as well as Analysis Order per Frequency. Again, it seems preferable to first upscale to third order using ab Image Upscaler. This plugin appears to work better than compass_decoder, but there is still some FFT flutter, and also choppiness when soloing individual channels, and there is a certain gravity towards individual speakers that somewhat reduces the perceived continuity of the sound field. The latter can probably be compensated for by adjusting plugin parameters.
Sparta ambiDEC decoding.
The sparta_ambiDEC plug-in employs a dual-band decoding approach, with options among several decoders for both registers: Sampling Ambisonic Decoder (SAD), Mode-Matching Decoder (MMD), Energy-Preserving Ambisonic Decoder (EPAD), and All-Round Ambisonic Decoder (AllRad, thus also covering what can be done with the IEM AllRAD decoder). These decoding algorithms work up to 10th order, so it makes sense to first upscale. Again, this is done using the AudioBrewer plugin. Upscaling to and decoding from 7th order works “too well”, giving too clear source separation, resulting in moving sound sources jumping from one speaker to the next. In the segment of the field recording that I use for testing, there is a train passing in front from left to right. At 7th order, independent of which decoder option I use, the illusion breaks down. Initially, the train stays in the Lss speaker, then jumps to the centre, and finally jumps to Rss. Reducing from 7th to 5th order gives a much better sense of continuity in the sound field.
Regardless of which decoding option is used, sparta_ambiDEC is the decoder that, so far, produces the most convincing results. It also sounds much more convincing when soloing individual channels.
ab Advanced Decoder
Decoding using ab Advanced Decoder.
The ab Advanced Decoder first performs internal upscaling to 7th order, using the same algorithm as the aforementioned ab Image Upscaler, before decoding with a beam-forming algorithm that, according to the developer, is especially optimised for a narrow spatial range with minimal side lobes.
Impressive as this plugin is, for this particular use case it ends up being “too good”. The upscaling to 7th order and subsequent beamforming work so well that the sound field starts to come apart. Rather than spatial continuity in the sound of the passing train, it jumps between speakers. There is a parameter to enable or disable the spatial upscaling, but this is a binary choice. If the plugin could be further enhanced to allow tuning the order of upscaling before decoding, so that I could go with 5th rather than 7th order, this could well be the preferred solution.
Upscaling and the Blue Ripple Rapture 3D decoder*
The final option tested is to first upscale to 5th order and then decode using the 7th order version of Blue Ripple’s Rapture 3D. For this part of the test, I experimented with two alternative upscalers: ab Image Upscaler and Penteo Pro+. Visually monitoring the resulting sound field, the difference between the two is quite informative.
Upscaling using Penteo Pro+, showing clear preference for placing the upscaled signal in the horizontal plane.
At its default settings, with AmbiX 1st order in and Ambix 7th order out, Penteo Pro+ produces a 7th-order sound field with clear emphasis of signals in the horizontal plane.
Even when maximising verticality and vertical diffusion, the resulting sound field gravitates towards the horizontal plane.
Even if the parameter for balancing the upper part is raised, the signal remains mostly in the horizontal plane. If vertical diffusion is added, the sound field starts spreading out, but even then, there is a clear preference for the horizontal dimension.
On other occations, the horizontal emphasis of Panteo Pro+ might be a welcome feature, but for the current material, it spatially alters the recorded sound sound field, in particular with respect to how the height speakers are used.
Upscaling using using ab Image Upscaler, showing a more even distribution over all of the sphere.
In contrast, ab Image Upscaler seems to give a more balanced and neutral upscaling, evenly maintaining the spatial distribution and spread in all directions while improving articulation and clarity.
I did not try additional upscaling options available through the Harpex algorithm or SPARTA up-scalers, as these had already been used in the previous decoding tests.
The solution I finally arrived at for this material is to upscale using ab Image Upscaler and then decode using Rapture 3D set to 5th-order decoding. Attempts at using 7th-order decoding yielded results similar to those of the ab Advanced Decoder, with too clear source separation resulting in a lack of continuity in the sound image. Using 5th-order decoding instead, I get results similar to what I imagine the ab Advanced Decoder could have produced, provided an added feature to set how far to upscale, rather than always upscaling to 7th order. The resulting decoding has a nice mix of spatial clarity and continuity, making it feel like a place with a continuous sound field. When soloing individual speakers, there is continuity in the sound, with little or no perceived artefacts of any kind.
1 Schwab, Michael. Transpositions: Aesthetico-Epistemic Operators in Artistic Research. Leuven University Press, 2018. https://www.jstor.org/content/oa_book_edited/j.ctv4s7k96


















