Preprint

Real-time spatial-audio vocoder reports higher spatial scores

Preprint: CSAVocoder reported higher spatial scores in baseline and listening comparisons, while audio-quality results were mixed.

A causal spatial-audio vocoder reported higher spatial-consistency scores than two listed spatialization comparisons when given ground-truth mel input, while its tested streaming configurations met the paper's real-time criterion. On the binaural test set, CSAVocoder reported ANG COS of 62.11% and DIS COS of 77.05%, with a real-time factor, or RTF, of 0.1587. Every reported streaming RTF was below 1.

The study asks whether multi-channel mel-spectrograms, maps of sound energy across time and frequency, and time-varying source-listener pose can be mapped to multi-channel spatial waveforms under strict causal streaming constraints.

A model built around spatial information

CSAVocoder combines a Spatial Adaptor, a spatial consistency discriminator and a strictly causal stateful generator. Its unified design supports binaural audio and four-channel First-Order Ambisonics, or FOA, output, which the study evaluates separately.

Training combines least-squares GAN, feature-matching, mel/STFT reconstruction and format-aware spatial losses. Binaural supervision uses inter-channel phase and level differences, known as IPD and ILD, while FOA supervision uses sound-field descriptors.

Testing the system

The training corpus contained roughly 600 hours and about 350k samples of binaural data, plus roughly 900 hours and about 310k samples of FOA data. All of it was sampled at 48 kHz. The authors randomly sampled 700 segments from all datasets for the test set and used a 9:1 train-to-validation split for the remaining material.

Baselines were evaluated by channel-wise inference for binaural and FOA outputs. The study combined objective measures of audio, spectral and temporal behavior, and spatial consistency with subjective listening tests.

The spatial measures led the comparison

On binaural output, the reported scores were 62.11% for ANG COS and 77.05% for DIS COS. The same evaluation reported MRSTFT 1.223, PESQ 2.109, MCD 2.153, periodicity 0.107 and RTF 0.1587, covering audio, spectral, temporal and processing measures.

With ground-truth mel input, CSAVocoder reported ANG/DIS COS of 62.11%/77.05%, compared with 26.65%/51.07% for GT+DSP and 50.83%/72.13% for GT+BinauralGrad. Its reported values were higher on both spatial measures in that comparison.

A mean-opinion-score, or MOS, study asked 29 participants to rate 200 randomly sampled test segments on a five-point scale. CSAVocoder received a spatial-perception score of 4.25 ± 0.16, the highest MOS-P among the models. Its quality score was 4.09 ± 0.21, below Vocos at 4.24 ± 0.11 and WaveFM at 4.17 ± 0.12.

A separate MUSHRA-style expert evaluation used nine experts and produced 120 dynamic results. CSAVocoder scored 83.0 for quality, 88.0 for spatial perception and 85.5 overall. It had the highest listed spatial-perception and overall scores, while its quality score was below Vocos and PriorGrad. The source reports intervals for these scores.

Comparisons between tested variants

In a direct comparison of two tested configurations, the causal CSAVocoder reported ANG/DIS COS of 62.11%/77.05%, compared with 60.21%/74.29% for the non-causal version. Its other listed measures were MRSTFT 1.223, PESQ 2.109, MCD 2.153 and periodicity 0.107, versus 1.083, 2.397, 1.742 and 0.093, respectively. These figures describe the tested variants rather than establishing why their scores differed.

An ablation comparison gave the four-head configuration ANG/DIS COS of 62.11%/77.05%. Configurations without the Mel Adaptor, spatial consistency discriminator, or Position Adaptor reported 42.60%/65.39%, 58.82%/74.63%, and 54.78%/70.63%, respectively. The pattern is descriptive of these tested configurations and does not establish that removing any one component caused the difference.

Streaming measurements were reported at chunk settings of 40, 60, 80 and 100. Mean compute latency was 15.24 ± 0.95, 15.15 ± 1.35, 15.52 ± 2.06 and 15.86 ± 2.40 milliseconds per chunk, respectively. The corresponding RTFs were 0.3811, 0.2526, 0.1941 and 0.1587, all below 1.

Beyond the main binaural test

The separate FOA evaluation produced Corr_all of 18.53, AUC_j_all of 63.44, MRSTFT of 1.248, MCD of 3.449 dB and PESQ of 1.972. AUC_j_all was the highest listed FOA value, while the audio-quality measures were mixed.

With a common low-pass cutoff of 7.8 kHz, CSAVocoder reported ANG/DIS COS of 68.71%/83.86%, compared with 68.64%/83.46% for Vocos and 66.05%/82.84% for WaveFM.

In pose-robustness tests, the unperturbed condition produced ANG/DIS COS of 62.11%/77.05%. At Gaussian noise level 0.2, the scores were 59.11%/74.83%; with 30% random pose-entry dropping, they were 60.18%/75.82%.

The test's boundaries

The evidence is limited to the preprint's computational experiments on recorded and simulated spatial speech and audio data, objective proxy metrics and small listening tests. The evaluations cover binaural and FOA, not higher-order ambisonics, loudspeaker layouts, object-based audio or personalized HRTF rendering. The listening panels were limited, and the main spatial measures were model-derived or representation-derived proxies rather than direct deployment-level localization outcomes.

CSAVocoder is described in an arXiv preprint, version 1, dated 26 Aug. 2026. The supplied evidence therefore speaks to the reported datasets, metrics, hardware measurements and listening studies, not to performance on other hardware, broad human deployment or generalization across all languages, accents, environments and pose conditions.

Paper data and sources

Original title: CSAVocoder: A Causal Spatial Audio Vocoder Towards Real-Time Spatial Audio Generation
Authors: Zhiyuan Zhu, Han Wang, Wenxiang Guo et al.
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-26
DOI: Not available
Original paper · Full text

Versions and corrections

  1. Published automatically after legal-source, freshness, evidence, and independent-verification gates passed.