Preprint

Audio model tracks theory-based rankings of musical dissonance

This preprint reports a computer-based representation that closely matched interval rankings but showed weaker, uneven results for some chords and modes.

The clearest result in the analysis is easy to state: Dissonance Spectrum, or DS, closely followed the study’s prescribed ranking of musical intervals. The minor second produced the largest summary value, while the octave produced the smallest. Spearman’s rho was 0.95055 and Kendall’s tau-b was 0.84615, two rank measures that together indicate a close match between the calculated and reference orderings.

But this was a test of a representation’s output, not a listening experiment. The controlled validation used fixed MIDI-rendered note events, generally with a piano timbre, and compared the calculations with predefined ordinal references. No listener judgments were reported, so the result says how DS organized the rendered audio rather than how people necessarily hear the intervals.

A representation built around relationships

DS is meant to make one kind of relationship visible: how simultaneous frequencies relate to one another. The work frames that explicit relation as a way to address the gap between energy-based spectra and latent learned representations.

The method supports two reference choices. In its intrinsic form, it uses the target spectrum as the reference; in cross-DS variants, it can use a fixed tonic or instrument spectrum. These variants let the analysis compare different kinds of reference relationships.

The preprocessing also differs by purpose. Downstream comparisons used global-maximum normalization. In the controlled theory tests, target and reference spectra were normalized separately, and the output of the relation transform was compressed afterward. Those details matter because the reported numbers belong to a defined measurement setup.

Where the rankings held—and where they did not

On four representative chord classes—major, minor, suspended and diminished—the calculated values landed in the intended order. The authors treated that as an ordering check, not a meaningful significance test, because there were only four classes.

When the analysis moved to 13 extended chord voicings, the match was positive but imperfect. Spearman’s rho was 0.62637 and Kendall’s tau-b was 0.46154, with local reversals in the ordering. The result suggests a broad association with the predefined pattern, not a flawless chord-quality scale.

Chord connections showed a similar split between a broad pattern and important exceptions. The ordinal code assigned 1 to tonic-function chords, 2 to predominant chords and 3 to dominant-function chords. It defined a connection through the first chord’s relation to C major, rather than through voice leading or learned temporal expectation. The resulting association was Spearman’s rho = 0.79373 and Kendall’s tau-b = 0.65465, although V7 and vii° differed substantially.

The seven church modes produced one of the stronger matches. Dsum, the summed dissonance summary, had Spearman’s rho = 0.89286 and Kendall’s tau-b = 0.80952 against the predefined order. But the trend was not perfectly stepwise: Dorian fell below Lydian.

That caution extends to the scale material. The analysis covered all 33 available scale files, but unranked scales were presented as a quantitative consonance palette for comparing outputs, not as a universal preference scale.

A stability check and exploratory edges

A separate loudness sanity check found little movement in the DS peak across seven independently rendered C4–C♯4 dyads. The Theil–Sen estimate was β = −0.000, with a 95% confidence interval from −0.014 to 0.000. Derived elasticity was 1.000, relative standard deviation was 0.94%, and relative range was 2.49%. In this test, the representation’s peak was therefore comparatively stable across the amplitude conditions.

The remaining examples were exploratory. Across seven software instruments, the location of the maximum differed, but loudness, spectral centroid—a measure of where the frequencies are centered—and envelope, the sound’s shape over time, were not matched; the comparison was therefore descriptive rather than a causal ranking of timbres. A 24-tone microtonal sequence, spaced in 50-cent steps, was analyzed without an experiential ranking.

A narrower conclusion

The strongest case made by these results is for an interpretable complement to learned music representations. DS exposes a defined class of frequency relation, then lets the output be checked against explicit reference orders. That is a narrower claim than saying the system has captured musical perception itself.

The work describes a separately submitted archive containing the DS implementation, extraction and training configurations, dataset manifests and audio hashes, evaluation and table-regeneration scripts, seed-level predictions and environment lock files. The computational runs used one A100 GPU per Slurm job on Linux, with PyTorch code and cached CQT and DS extraction.

Taken together, the evidence supports a narrower conclusion than a universal preference scale. In controlled rendered examples, DS tracked several predefined orderings and showed a stable loudness response, while some chord and mode details reversed or deviated. Because the validation was computational rather than listener-based, it does not establish that people will perceive the reported interval, chord or mode rankings in the same way.

Paper data and sources

Original title: Dissonance Spectrum explicitly models perceptual frequency interactions for better music understanding
Authors: Tianle Wang, Xinyi Tong, Liangke Zhao et al.
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-26
DOI: Not available
Original paper · Full text

Versions and corrections

  1. Published automatically after legal-source, freshness, evidence, and independent-verification gates passed.