A new arXiv preprint reports that Stream4D video outputs scored better on reconstruction, motion preservation and overall preference than the corresponding distilled bases across three streaming autoregressive video backbones. The main evaluation used a 500-prompt testing set focused on prominent motion and included comparisons with World-R1 and VideoGPA.
The paper asks whether replacing a static 3D reconstruction critic with a dynamic 4D reward can address two problems in long rollouts: geometric drift and loss of motion. Here, 4D reconstruction refers to judging scene consistency as the generated video changes over time, rather than treating the scene as static.
A reward built around motion and change
Stream4D evaluates each rollout with a feed-forward 4D reconstruction score, a gated motion-quality term and a lightweight perceptual anchor. It puts those reward axes on a common standardized scale through per-axis z-normalization, then applies forward-process DiffusionNFT optimization.
Training left the base models frozen and used LoRA adaptation with the forward-process objective. The evaluation covered native short windows for Self-Forcing and Causal-Forcing, along with a longer LongLive window.
Reported reconstruction gains were largest on LongLive
On the paper’s main 4D-PSNR measure, Stream4D’s reported gains were 3.46 dB for Self-Forcing, 5.53 dB for Causal-Forcing and 6.76 dB for LongLive. SSIM and LPIPS were also reported as better on all three backbones.
A second evaluation, 4DGT, uses architecture, weights and training data disjoint from MoVieS. Stream4D ranked best on PSNR and SSIM in every backbone block and led World-R1 by 0.7, 1.1 and 2.5 dB in the same order. But 4DGT shares the StreamVGGT camera estimator with the other reconstruction evaluation, and the paper cautions that the result does not certify metric-accurate geometry.
Learned judges also favored the method
A vision-LLM judge reported Stream4D consistency scores of 82.2%, 73.9% and 74.2% for Self-Forcing, Causal-Forcing and LongLive, compared with 75.9%, 69.1% and 54.0% for World-R1. Its motion-preservation scores for Stream4D were 0.83, 0.77 and 0.71, respectively.
VideoReward, another learned video-quality model, gave Stream4D overall win rates against the corresponding base of 66.2%, 76.0% and 84.4%. World-R1’s reported values were 61.8%, 74.8% and 78.2%; Stream4D’s motion-quality margins over World-R1 were 12.2, 6.4 and 11.4 percentage points.
The human test was small but pointed in the same direction
In the human pairwise study, Stream4D won overall against the base 60% of the time, against World-R1 76% and against VideoGPA 80%. Its motion win rates were 43%, 72% and 87%, respectively.
The test used 50 high-motion prompts, produced 150 comparisons and involved five raters. The human evidence was therefore limited in scope compared with the broader model-based evaluation.
The results came from a balancing act
Reward ablations reported that no individual axis could be removed across the evaluated backbones. Removing the motion term was associated with motion collapse; removing reconstruction was associated with high motion but a collapse in coherence; and removing the perceptual anchor was associated with losses in consistency or preference on some backbones.
On a uniformly random 500-prompt subset, Stream4D was the only method reported to beat its base under the joint verdict on all three backbones. Its Self-Forcing motion score was 0.816 and its LongLive consistency score was 68.5%.
Important limits remain
The comparisons show reported differences between systems, but they do not establish that Stream4D caused those differences. The analysis reports no confidence intervals, p-values or formal significance tests.
Two measurement caveats are especially important: the vision-LLM results come from a single learned evaluator, and the 4DGT check still shares the StreamVGGT camera component with the other reconstruction evaluation. The main test also centered on motion-prominent prompts.
The authors interpret the findings as evidence that the recipe transfers across the evaluated distilled backbones while preserving motion and preference quality. They propose a streaming 4D reconstructor and action- or camera-conditioned grading as next steps.
The work is an arXiv preprint posted on 20 August 2026.
Paper data and sources
Original title: Stream4D: 4D-Consistency for Streaming Autoregressive Diffusion Video Models
Authors: Yuanhao Ban, Jiaqi Feng, Hengguang Zhou et al.
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-20
DOI: Not available
Original paper · Full text