An arXiv preprint reports that its AvatarDynamizer method posted the best image- and video-level generative metrics among the displayed methods, while ranking second on reconstruction quality. The system is designed to add pose-dependent appearance and deformation to off-the-shelf static 3D human avatars and render them consistently from multiple views.
The paper asks whether faithful, pose-dependent surface dynamics can be added to static avatars without losing control over how they appear from different viewpoints.
How the system adds motion
Its three-step pipeline encodes multi-view videos as 2D dynamic textures and, together with pose-related normal maps, decodes them into 3D Gaussian splats. The method is intended to turn static avatars from diverse input modalities into controllable 4D avatars.
The evaluation used PSNR, SSIM and LPIPS for image reconstruction, alongside FID and FVD for image- and video-level generative quality.
The preprint is version 1, dated 20 August 2026. Generalized Gaussian Decoder training used 56 DynaHuman subjects and 478 MVHPP subjects, while benchmark scores were averaged over two subjects per dataset.
The reported benchmark edge
In the main comparison, the displayed Ours entries recorded FID/FVD pairs of 22.58/56.72 and 15.28/37.00. The authors describe these as the best generative results among the displayed methods, while reconstruction quality ranked second.
A representation test favored dynamic reference textures over static ones when both used the decoder: PSNR was 32.02 versus 30.36, SSIM was 0.874 versus 0.808, and LPIPS was 0.103 versus 0.160.
On DynaHuman testing, joint training reported PSNR 29.71, LPIPS 0.122 and FVD 22.97. Excluding DynaHuman produced 28.77, 0.142 and 33.60, while excluding MVHPP produced 29.70, 0.124 and 25.77 on the same reported measures.
A separate dynamics-prior test reported PSNR 26.61, SSIM 0.745, LPIPS 0.182, FID 13.84 and FVD 37.74 for the full system row. Its caption says that removing pretrained diffusion weights or temporal information degraded performance, although the extracted table has misaligned labels and columns for some rows.
These comparisons were reported without uncertainty intervals or significance tests.
A small preference test
In a separate user study, 18 participants compared results for five subjects. The authors report that the proposed configuration was preferred for identity preservation, physical plausibility and overall quality, and ranked second for view consistency.
On five-point ratings, the proposed configuration scored 4.000 for identity consistency, 4.003 for physical plausibility, 3.789 for view consistency and 3.947 overall.
In another comparison limited to two reported DynaHuman subjects, LHM plus Ours had lower FID and FVD than Wan-Animate 2 on both subjects.
Where the method may struggle
The method may be less robust for garments with extreme non-rigid deformation. Its fixed SMPL-X template limits explicit modeling of topological clothing changes, while self-occlusion can leave texture observations incomplete.
The findings therefore describe comparative performance on rendered-avatar benchmarks and a separate preference study, rather than a causal test or evidence of real-world deployment safety.
Paper data and sources
Original title: AvatarDynamizer: From Static to Dynamic Human Avatars via Generative Dynamic Textures
Authors: Guoxing Sun, Heming Zhu, Linjie Lyu et al.
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-20
DOI: Not available
Original paper · Full text