A preprint describes Luce, a model that takes a single image and generates a relightable 3D asset. Its central design choice is to keep geometry and PBR materials—the surface-property data used for physically based rendering—in one voxelized, multimodal Gaussian cloud. A variational autoencoder compresses that representation, and a rectified-flow transformer generates the latent from image features drawn from multiple layers. In plain terms, Luce tries to represent an object's shape and material together rather than treating them as separate outputs.
The output is not limited to one 3D format. The generated latent can decode into relightable PBR Gaussians, and Luce can optionally produce a textured mesh with a tangent-space normal map. The design therefore offers one image-conditioned route to a Gaussian representation and another to a textured mesh, with both intended for use as 3D assets.
The tests behind the scores
Training used approximately 500,000 PBR-filtered assets from Objaverse and Objaverse-XL, alongside a 158,000-asset PBR subset of TexVerse. Evaluation covered a 338-asset PBR Toys4K reconstruction subset, 412 Toys4K generation assets and 130 AI-generated images. The comparisons included TRELLIS, TRELLIS 2, LiTo and 3DTopia-XL.
The clearest numerical result came from Toys4K generation. FID, one of the paper's scores for comparing generated results, was 20.99 for Luce GS, the Gaussian-output version, versus 29.22 for TRELLIS 2. Luce GS had the lowest reported FID among the compared methods, with the paper describing the gap as more than 8 FID. No uncertainty estimate was reported for this comparison.
Strong results, but not a clean sweep
On the AI-generated-image benchmark, Luce GS also led the reported CLIP and SigLIP2 comparisons. Its scores were 0.8519 on CLIP and 0.8508 on SigLIP2, compared with 0.8299 and 0.8339 for TRELLIS GS. The Luce mesh variants led the mesh-only ULIP and Uni3D-L comparisons. In the paper's setup, these measures were used to judge input-output alignment.
Reconstruction results were more mixed. Luce GS had the best reported color PSNR at 36.1 dB and normal PSNR at 34.6 dB. TRELLIS 2, however, led the albedo and metallic-roughness measures; for albedo, Luce's LPIPS was 0.055 versus 0.051 for TRELLIS 2. The Gaussian version was therefore not ahead on every reported reconstruction measure.
The authors also tested the image-conditioning design in an ablation. The multi-layer variant recorded a Toys4K FID of 20.99, versus 25.21 for a single-layer variant. On the AI benchmark, its CLIP score was 0.8519 versus 0.8081. The reported comparison favored the multi-layer setup on both measures.
For mesh output, the tangent-normal variant reached a normal PSNR of 33.0 dB, compared with 29.5 dB without tangent transfer. Those figures come from the paper's mesh comparison and concern the normal-map part of reconstruction.
The boundaries of the result
Speed and model size varied by output path. On a single H100, mean inference took about 42 seconds for the Gaussian path. The mesh path took about 148 seconds without the tangent normal map and about 159 seconds with it. The reported parameter counts were 4.5 billion for the Gaussian path and 4.6 billion for the mesh path. Only mean timing was reported, with no variability estimate.
The authors interpret the design as delivering state-of-the-art generation quality while producing relightable outputs by design. They also identify important constraints: finite voxel resolution can leave features spanning only a few voxels under-resolved, and the current material model does not explicitly cover complex appearance effects. Those limitations sit alongside the benchmark results and narrow what the reported scores establish.
Paper data and sources
Original title: Luce: Relightable Gaussians for 3D Asset Generation
Authors: Mayank Singh, Michele Stoppa, Alvise Memo et al.
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-25
DOI: Not available
Original paper · Full text