A frequency-aware stereo autoencoder had the best point estimates on five of seven reconstruction metrics in a test of 546 tracks. The comparison included four recent open-source VAEs, with εar-VAE2 evaluated both without and with its Duplex-Aware Refiner.
The paper is an arXiv version 1 preprint dated 20 August 2026. Its central question was which modeling choices best preserve perceptually salient spectral, phase and stereo information at a fixed compression rate.
How the model handles sound
The system, εar-VAE2, is an autoencoder built to compress and reconstruct sound. It uses complex-spectrum stereo separation and a frequency-aware nonlinearity, then encodes 48 kHz stereo music into continuous latent representations—compact numerical summaries used for reconstruction—at 25 Hz, with 128 dimensions and 1,920-fold temporal downsampling.
Spec-SnakeBeta, the named activation function, learns one pair of parameters for each frequency bin and shares that pair across feature channels, rather than using channel-wise or fully independent channel-frequency parameters.
The optional Duplex-Aware Refiner applies different corrections across the spectrum: phase-only below 1.5 kHz, joint magnitude-and-phase correction from 1.5 to 4 kHz, and magnitude-only correction above 4 kHz.
Mixed results across representations
To isolate representation choice, the study compared five compact autoencoders at approximately 3.0 million to 3.1 million parameters, matched within ±5%. All used a 128-dimensional latent at 25 Hz and trained for 100,000 steps.
Complex STFT had the lowest full-band spectral distance, a measure of mismatch in frequency content, at 1.122, and the lowest high-frequency distance at 1.178. The high-frequency figure was 22% lower than the waveform-patch result, but waveform patch led on three of four aggregate metrics and on low-frequency distance.
In the broader reconstruction comparison, the complete εar-VAE2 system had the best point estimates on STFT, log-scaled STFT, Mel, log-scaled Mel and spectral-projection error, and tied the best cross-channel phase-consistency score.
Adding the refiner corresponded to a lower Mel Distance: 0.461 with it versus 0.572 without it, a reported 19.4% reduction.
A narrower refiner scored better in the test
The banded refiner also compared favorably with the Unconstrained Refiner in the reported point estimates. It used 532 residual dimensions per channel instead of 962, about 45% fewer, and had lower Mel Distance (0.658 versus 0.705), lower STFT Distance (1.019 versus 1.120), and lower errors for low-frequency inter-channel phase difference (IPD), low-frequency inter-channel level difference (ILD) and high-frequency ILD.
Ten professional mixing and mastering engineers gave the banded design a higher mean paired rating than the Unconstrained Refiner, 0.75 versus 0.66.
Generation scores favored the same system
In downstream generation tests, each system produced 100 English songs and 100 Chinese songs. εar-VAE2 had higher point estimates than LeVo 2 on all 12 reported automatic metrics, while the refined εar-VAE2 row was higher again than the unrefined row on every listed metric.
In an activation ablation using three seeds, the F-Log activation had the best mean on four of six metrics. Its SI-SDR score, one measure of signal fidelity, was 4.40 dB, versus 3.46 dB for F-Uniform, and it used 2,730 activation parameters instead of CF-Log's 350,720—about 128 times fewer.
What the scores leave unresolved
The paper reports these comparisons mainly as point estimates, without confidence intervals or significance tests. It describes the controlled studies as establishing relative rankings rather than production-scale limits.
The main cross-system reconstruction comparison used native model configurations rather than matched budgets, and the banded and Unconstrained Refiner variants did not have matched output capacity.
The listening test involved 10 engineers, and automatic generation scores do not by themselves establish human-perceived superiority.
The full-scale model used approximately two million publicly sourced tracks plus about 10,000 hours of proprietary 48 kHz stereo music. The detailed source mix was withheld for data-strategy and licensing reasons.
The findings do not establish that complex STFT is uniformly better than waveform representations, or that higher automatic generation scores would translate into better human judgments.
Paper data and sources
Original title: Fourier is Frontier: Frequency-Aware Autoencoding for High-Fidelity Music Reconstruction
Authors: Kangdi Wang, Yusheng Dai, Jin Xu
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-20
DOI: Not available
Original paper · Full text