Automatic scores used to judge whether generated video and audio line up may be measuring different parts of the problem, according to a methods audit. Four metrics often disagreed, and the metric that best tracked an imposed timing shift was not the one that most closely matched the study’s human-aligned proxy.
The paper, posted as a preprint and accepted for a non-archival poster presentation at the ECCV 2026 Workshop, does not identify a single winner. Instead, it points to audio-visual synchronization as a set of separate abilities, with results reported by disruption type and with some indication of ranking uncertainty.
A test with known disruptions
The researchers audited AV-Align, ImageBind AV-relevance, JavisScore and Synchformer/DeSync under one reliability protocol. They tested whether scores changed consistently as clips became less synchronized, whether preprocessing choices altered results, whether the metrics agreed with one another, and whether they tracked PEAVS, an automatic proxy used to represent human-aligned judgments.
For the main audit, they used 75 real, well-synchronized clips, with five clips in each class. They then imposed controlled, increasingly severe distortions: shifting the timing, changing audio speed, shuffling audio fragments and adding intermittent mutes. A separate two-times replication used 150 clips.
The analysis used Kendall’s tau, a rank-correlation measure, to compare metric scores with the known severity of each distortion. The researchers also used clip-level bootstrap resampling with at least 1,000 resamples, along with measures of preprocessing sensitivity and the chance that a small change would reverse the ranking of two clips. For a combined score, they tested a regularized ridge model in strictly out-of-fold five-fold and leave-one-out evaluations.
Strong on timing, weaker on other signals
The clearest split appeared in the synthetic tests. Synchformer/DeSync tracked temporal shifts best, with a Kendall’s tau of 0.84. But AV-Align was weak at following audio-speed changes and shuffled fragments, with values of 0.04 and 0.01. ImageBind and JavisScore performed best against the PEAVS proxy, tying at 0.20.
That pattern held in a separate set of 75 VGGSound clips from less controlled material. Synchformer reached a temporal-shift tau of 0.63, while the other metrics reached no more than 0.27. JavisScore, however, reached 0.59 for the mute test. The result suggests that a metric’s apparent strength can depend heavily on what kind of mismatch is introduced.
The scores were also vulnerable to how clips were prepared. AV-Align had a median coefficient of variation, a measure of relative score instability, of 0.54, compared with no more than 0.15 for the other metrics. In the distortion family where each metric was least stable, the chance of an adjacent ranking flip was 0.96 for AV-Align, 0.69 for ImageBind, 0.90 for JavisScore and 1.00 for Synchformer.
Combining scores did not solve the problem
Across the audit, agreement among the metrics was very low, with a Krippendorff alpha of 0.066. The main exception was the ImageBind-relevance and JavisScore pair, whose rankings had a Kendall’s tau of 0.78. Against PEAVS, Synchformer had a tau of 0.07, ImageBind and JavisScore tied at 0.20, and AV-Align was approximately zero.
The researchers also tested whether a learned combination could produce a more human-aligned result than any single metric. It did not in this analysis. Held-out ridge fusion reached no more than 0.12 against PEAVS, below the best individual value of 0.20, while leave-one-out k-nearest-neighbor fusion reached no more than zero.
For generated audio-video outputs, the metrics were relatively consistent when the gap between two results was large: the probability of a far-pair ranking flip was no more than 0.09 for every metric. For a close pair, the flip probabilities were 0.35 for AV-Align, 0.33 for ImageBind-relevance, 0.35 for JavisScore and 0.08 for Synchformer/DeSync.
A warning against one-number evaluations
The findings point toward reporting a reliability card rather than one bare synchronization number. Such a report would show how a metric behaves across different disruption families, include uncertainty, and make clear how confidently it can distinguish nearby results. The audit supports that approach because the metrics’ strengths did not line up on a single scale.
The findings do not show that Synchformer/DeSync is perceptually best overall, or that any metric works reliably across every kind of distortion. They also do not establish direct agreement with human preferences. PEAVS was an automatic learned proxy rather than fresh human labeling, and its representation may share biases with embedding-based metrics.
The evidence here is limited to a reliability audit using controlled distortions, agreement tests and specific fusion methods. The fusion result does not rule out other approaches, including calibration against direct human labels. The paper leaves open whether broader data, generators and real-world artifacts would change the ranking of the metrics.
Paper data and sources
Original title: What Do Audio-Visual Synchronization Metrics Actually Measure?
Authors: Jai Kumar Sharma, Peeyush Tapadiya
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-25
DOI: Not available
Original paper · Full text