Preprint

New benchmark tests whether AI video scores match human judgment

Preprint: An expanded benchmark uses more than 60,000 generated videos and 36,000 new annotations to assess aesthetics and generation quality.

A new arXiv preprint describes a larger, more structured way to judge AI-generated video, combining expanded human annotations with automated systems for visual appeal, tags and generation quality. The benchmark adds 36,000 task-level annotations, and its evaluators are tested against held-out human ground truth. In a separate experiment, one evaluator score was used to fine-tune Wan2.1; the reported average aesthetic score rose from 0.49 to 0.52 on a held-out test set.

That increase is the benchmark evaluator's score, not a direct vote from viewers. The authors report qualitatively more appealing visual composition after fine-tuning, but the supplied analysis identifies no independent post-optimization human-preference measure. The result therefore shows an evaluator-score increase, without establishing that people would prefer the resulting videos.

A bigger testing ground

VGA-BenchV2 retains a fine-grained taxonomy of 52 sub-dimensions covering aesthetics and generation. It draws on 1,016 prompts and more than 60,000 videos generated by 12 mainstream models.

The expansion adds 16,200 annotations for aesthetic quality, 13,200 for aesthetic tagging and 6,600 for generation quality. Compared with VGA-Bench, the paper reports scale-ups of 13.46 times, 11.15 times and 1.55 times in those three areas, respectively.

The annotation protocol used expert exemplars and guidelines, trained multiple annotators and batch audits. Work could be rejected and sent for re-annotation. Aesthetic ratings used a 0-to-10 scale and were averaged across independent ratings; tags were decided by majority vote. Generation quality was captured with structured questions using ordered answer choices.

Different tools for different judgments

The evaluation system is split by job. VAQA-Net performs specialized aesthetic regression, which predicts a numerical quality score. VTag-Net and VGQA-Net use a large vision-language model for semantic reasoning in tagging and generation-quality judgments.

In validation against held-out human ground truth, VAQA-Net reached 87.6% on Overall Score SROCC. SROCC is a rank-based comparison, so the figure reflects how closely the evaluator ordered videos to match the human labels. VTag-Net's accuracy was 93.6% for Light Color and 83.7% for Number of Light Sources.

VGQA-Net averaged approximately 71.3% accuracy across 31 generation-quality dimensions. Its reported accuracy reached 89.2% for Scene Realism and 85.7% for Abnormal Lighting Detection. The spread suggests that agreement with the labels differed by dimension, even within the same evaluator.

Scores, rankings and a narrow optimization test

For model comparisons, the researchers used videos held out from evaluator training to prevent data leakage. Aesthetic quality was represented by the average predicted score, tags by alignment accuracy and generation quality by the mean across sub-dimensions.

Among the reported aggregate results, Sora2 had the highest aesthetic score at 0.50 and the highest generation-level score at 0.80. Mochi had the highest tag-classification value at 0.68. Different leaders appeared for different metrics.

Those rankings are benchmark outputs, not proof of causal superiority. The comparison was nonrandomized, no conventional control arm was reported, and the supplied analysis includes no uncertainty estimates or inferential comparisons between models.

The optimization test was narrower. VAQA-Net's Overall Score served as the reward for fine-tuning Wan2.1 with Flow-GRPO, LoRA and ODE-to-SDE conversion. On the held-out test set, the average aesthetic score moved from 0.49 to 0.52. The reported result is evidence of a score change in one experiment, not a broader demonstration of human preference.

The supporting record

The paper states that benchmark resources are available online. It also says the benchmark content was screened for offensive material and that annotation procedures received Institutional Review Board approval.

Paper data and sources

Original title: VGA-BenchV2: An Expanded Unified Benchmark and Multi-Model Framework for Evaluating Video Aesthetics and Generation Quality
Authors: Longteng Jiang, DanDan Zheng, Qianqian Qiao et al.
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-26
DOI: Not available
Original paper · Full text

Versions and corrections

  1. Published automatically after legal-source, freshness, evidence, and independent-verification gates passed.