Preprint

Preprint uses sound spectra to measure timbral distance

An arXiv preprint compares vocal and instrumental spectra, then explores whether the same geometry can inform fingerings and chord voicings.

A distance between two voices

An arXiv preprint reports a Wasserstein distance, a measure of how far apart two frequency distributions are, of approximately 442 Hz between chest voice and falsetto in an illustrative comparison. The paper treats each normalized instrumental-timbre or speech spectrum as a probability density for perceived frequency. In the same example, the spectral centers of mass were approximately 2,066 Hz for chest voice and 1,713 Hz for falsetto, a reported shift of 353 Hz and residual timbre change of about 89 Hz. The values describe the model’s comparison of example sounds, not a listener score.

That comparison separates a shift in the average spectral location from the rest of the measured difference. The paper states a general bound: for two spectra, the Wasserstein distance is at least the absolute difference between their centers of mass. In plain language, the overall distance cannot be smaller than the gap between the spectra’s average frequencies. In this example, the approximate 442 Hz distance is larger than the 353 Hz center-of-mass shift, while about 89 Hz is reported as residual timbre change. The values are approximate point estimates, and no uncertainty interval or replication count is reported.

A map built from examples

To create the inputs, the paper applies a short-time Fourier transform, a frequency analysis performed over brief portions of a sound. This is a methods study built around illustrative vocal and instrumental spectra, not a participant cohort. Its examples use sounds recorded across the lowest-to-highest register of a single instrument, and the document does not report a total sample count. For flute and viola spectral sequences, it constructs Wasserstein self-distance matrices, showing the modelled distances within each sequence. It also builds an expanded mutual-distance matrix from timbres formed through direct products, extending the comparison across the sequences.

A separate mapping focuses on the clarinet. Using persistent homology, a topological mapping method, the paper preserves Wasserstein distances, clusters nearby clarinet timbres by register and places the spectrum near Throat G in a distinct distribution. The separation is reported qualitatively, with no cluster metric, inferential test or uncertainty estimate given for the mapping.

When the map reaches the instrument

The paper also applies the geometry to fingering choices. In collaboration with an oboist, the authors devised 13 new F fingerings, compared their spectra with clarinet Throat G and reported a fingering method with notably high compatibility. In the bassoon comparison, Professor Turnovsky’s fingering yielded a small Wasserstein distance whether the clarinet’s Throat G or an alternative clarinet fingering was used.

These examples show the kind of cross-instrument matching the authors are trying to quantify. They do not establish that either fingering improves ensemble blend or performance. The oboe result has no numerical compatibility value in the reported text, while the bassoon comparison supplies no numerical distance, uncertainty estimate or test of the outcome in ensemble performance.

From chord shape to harmonic motion

For a three-voice chord configuration, the paper proposes the area of a triangle whose sides are the three pairwise Wasserstein distances as an index of timbral affinity or dissociation. It invokes Heron’s formula to calculate the area. The same three distances also feed a second-moment measure, defined as the mean of their squared values, and a related Wasserstein Timbre Deviation intended to quantify timbral dispersion. These measures are designed to describe how a group of three sounds is arranged, not only how two sounds differ.

The broader aim is to extend Schönberg’s question theoretically from the perspectives of composers and players. The authors propose timbral-harmonic progressions that can be divergent, convergent, parallel to or contrary to conventional cadential motion. They also state that the timbral-harmony model approaches classical harmony as the Wasserstein index tends to zero. The limiting relation is presented as a theoretical proposal and has not been empirically evaluated in the supplied analysis.

A framework still waiting for a listening test

The central open question is whether a physical spectral distance corresponds to what listeners hear as timbral similarity, affinity or blend. The framework assigns normalized spectra the role of probability densities for perceived frequency, but that interpretation has not been validated with direct listener data. Nor has the fingering analysis been tested by measuring actual ensemble blend or performance quality.

Another constraint is the model’s reach beyond three voices. The paper says that a general area-or-volume evaluation is not generally available for four or more dimensions because Wasserstein distances may fail the required Gram-matrix conditions, the checks needed for such a geometric construction. The work remains an arXiv preprint, and no funding source is reported in the supplied document.

Paper data and sources

Original title: Klangfarbenakkord and Klangfarbenharmonien Metric Space Models for Music on Informational Geometry 1
Authors: Yusei Tamura, Shigekazu Ishihara, Ken Ito
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-28
DOI: Not available
Original paper · Full text

Versions and corrections

  1. Published automatically after legal-source, freshness, evidence, and independent-verification gates passed.