Preprint

Audio models falter when sounds and listeners change

A preprint analysis finds target labels recover only part of the loss, while physiological and individual-level prediction remain much harder.

Audio models can perform well within one sound domain but lose most of that performance when moved between music and environmental recordings, a preprint analysis finds. In transfers among corpora from the same sound domain, the mean Spearman rank correlation was 0.592 across 36 ordered pairs. That statistic measures whether the model preserves the ordering of ratings. Across 36 transfers between music and environmental sound, the mean was 0.054.

The work examines how much performance survives four boundaries: new material, edited audio, physiological response and individual listeners. It is a secondary analysis with no new data generated. The researchers used four rated-sound corpora, four pretrained audio representations and three physiological-recording corpora. Six regression algorithms were tested with five-fold cross-validation grouped by source recording. For transfer tests, models were fitted on one corpus and evaluated on another without refitting.

A strong in-domain score did not travel

The results show why within-corpus scores can be misleading. On the 717-excerpt PMEmo corpus, the model's correlation with observed ratings was 0.817, equal to 84.4% of the annotation-reliability ceiling. The 95% confidence interval for that share was 81.5% to 87.0%. In plain terms, the model came close to the agreement limit built into the ratings on that corpus, but that result did not carry over to other sound domains.

Adding target-side labels recovered some of the lost ground. With 10, 25, 50, 100 and 200 labels from the target corpus, the analysis recovered 20%, 34%, 50%, 64% and 69% of the corpus-swap gap. Gains slowed sharply between 100 and 200 labels, and the tested budgets did not close the remaining gap.

Changing the model's input representation did not generally solve the transfer problem either. For the three uncontaminated representation families, corpus swaps cost 20% to 31% of within-domain performance and 84% to 95% across domains. The comparison has a narrow base: CLAP was handled separately because its pretraining data overlapped the environmental corpus, and only one environmental corpus was available.

One representation broke the pattern

One result did challenge the broader pattern. In a predeclared test, adding 39 tonal descriptors produced a 0.132 gain in happiness prediction, compared with a mean gain of 0.045 on the other seven axes. Tonal descriptors alone reached a correlation of 0.580, versus 0.474 for 122 spectro-temporal descriptors. The result concerns model predictions, not listener responses, but the authors present it as evidence that a representation change can selectively improve one target.

Synthetic audio edits produced another model-level signal. Across 80 source excerpts, spectral smoothing led to lower predicted arousal in both music and environmental sound. The median slopes were 0.015 and 0.074, respectively, with the same sign in 70% and 71% of excerpts. Only spectral smoothing showed reliable excerpt-level consistency, and the study did not measure listener responses, so the finding describes model behaviour rather than a confirmed human reaction.

Physiological signals were weaker

The physiological boundary began with exclusions. Before substantive analysis, two of the three physiological corpora were removed: an EEG stimulus-code and audio mapping could not be reproduced, with a permutation p-value of 0.19, and none of 94 electrodermal participants passed both predeclared positive controls.

In the EEG analysis, the strongest intraclass correlation, a measure of consistency across trials, was 0.092 across 1,240 trials and 156 EEG measures. The best-agreeing self-report scale reached 0.221 on the same trials. The 95% intervals were 0.037 to 0.142 for the EEG maximum and 0.155 to 0.281 for self-report. Because the EEG value was the maximum found among 156 measures, the paper treats it as a bound from that search, not as evidence that EEG contains no information.

Across all tested information sources, none exceeded 31% of the physiological ceiling. The strongest pairing was subjective valence rating against skin-conductance response rate, with a correlation of 0.117 against a ceiling of 0.382, or 30.6% of it. Acoustics-predicted ratings reached 24.5%, while the leading acoustic principal component reached 21.3%. The study did not measure a physiological target-side learning curve, so these low results do not establish that a larger data budget could never help.

Individual prediction demands repeated observations

For individual listeners, the numbers grew quickly as the expected association weakened. Simulations targeting 80% power estimated 314 observations per person for an association of 0.40, 564 for 0.30, 1,364 for 0.20 and 4,888 for 0.10. These were within-person associations, meaning the model and a given listener would need to line up repeatedly for a detectable relationship. The comparison corpus contained 75 listeners. These are simulation-based estimates, not observed thresholds.

A separate calibration analysis estimated 67 observations per person for urban soundscape ratings, 166 for music ratings and 318 for music electrodermal responses. At 42 observations, two of eight soundscape cells crossed the study's break-even point, while 11 of 15 electrodermal cells were indistinguishable from their null. The authors describe the two soundscape crossings as candidate signals rather than validated results.

A map of the remaining uncertainty

Taken together, the paper offers a diagnostic split. Some gaps appear partly recoverable with target-side observations, as the label curve was; others were not recovered by the tested representations or information sources. That leaves open whether more target-side physiological data would recover the gap, whether the cross-domain barrier persists with an environmental corpus outside the representations' pretraining data, and whether model responses to edited audio predict actual listener responses.

The document is an arXiv version 1 preprint dated 27 August 2026. The analyses used public data, and derived tables, code, run records and figure scripts were made available. The authors note limits including uncommitted working trees and incomplete or absent hashes, so exact bit-for-bit reproducibility was not established. Funding was not reported. The authors did report involvement in developing a commercial sleep-audio product.

Paper data and sources

Original title: Not all generalisation failures can be bought back: four boundaries in affective audio modelling
Authors: Jingyi Zhang, Xiaotong Yao
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-27
DOI: Not available
Original paper · Full text

Versions and corrections

  1. Published automatically after legal-source, freshness, evidence, and independent-verification gates passed.