Preprint

MTAR Reports Faster Training and Lower FID Than LlamaGen

Preprint: ImageNet tests report up to 0.95 lower FID and 39% faster training for MTAR than LlamaGen, with no change to inference.

A training-stage change

An arXiv preprint reports up to 0.95 lower FID and 39% faster training for MTAR than LlamaGen. FID is a benchmark score for generated images in which lower values are preferred. The proposed components are used only during training, and the report says they add no overhead or modification to autoregressive inference.

MTAR combines multi-token prediction (MTP), token-level contrastive regularization (TCR) and semantic dropping (SD). The paper presents them as a unified framework aimed at denser supervision, more discriminative token representations and greater training efficiency. MTP, TCR and SD are applied only during training.

The benchmark comparisons

The tests focused on class-conditional ImageNet generation at 256 × 256 resolution. The ablations used 100k training images and evaluated 10k generated images against 10k real validation images. The full-scale assessment generated 50k images.

On the full benchmark, MTAR-B reported FID 4.50, IS 213.96, precision 0.84 and recall 0.48. MTAR-L reported FID 2.85, IS 271.41, precision 0.83 and recall 0.54. FID favors lower values, while IS, precision and recall favor higher values in the reported scoring.

Training-resource comparisons also favored MTAR in the reported figures. For the Base models, MTAR training took 30 hours and used 31 GiB of memory, compared with 115 hours and 45 GiB for LlamaGen. For the Large models, the corresponding figures were 42 hours and 73 GiB for MTAR, versus 175 hours and 97 GiB for LlamaGen.

At the same number of training iterations, MTAR-B and MTAR-L showed reported FID differences of 0.96 and 0.95 in their favor, with speedups of 1.27× and 1.39×. Under smaller training budgets, the report says MTAR-B surpassed LlamaGen-B at about one-third of the iterations, while MTAR-L outperformed its baseline at about one-sixth. The corresponding reported speedups were 3.81× and 8.33×.

The component comparisons

In the component ablation, the baseline had FID 21.96 and IS 54.04. The version with MTP corresponded to FID 19.75 and IS 62.12, the largest individual reported quality change. TCR's reported FID was 20.55, while SD alone had FID 22.02 and a 1.62× speedup. The combined MTP+TCR+SD configuration had FID 18.65 and a 1.27× speedup.

Within the tested MTP variants, MTP (R1, B) had FID 19.75, compared with 20.13 for MTP (R1, R2) and 19.96 for MTP (R1, B.R1), with one auxiliary head. A larger-head configuration had FID 20.95. The reported numbers did not show a consistent benefit from adding heads.

For TCR, the lowest listed FID among the tested sample sizes came from 2,048 sampled tokens, at 18.78. The reported FID was 19.83 with 8,192 tokens and 20.70 with 16,384. The listed results did not favor simply sampling more tokens in the settings tested.

The semantic-dropping comparison included both encoder choices and training schedules. DINOv3-guided dropping had FID 18.65, compared with 19.97 for random dropping and 19.70 for SigLIP2. A 50% drop rate was paired with FID 18.65 and a 1.60× speedup. The 60%:40% schedule had a slightly lower FID of 18.41, while the adopted 80%:20% schedule had FID 18.65 and the same reported 1.60× speedup. The authors describe the selected settings as a quality-efficiency compromise.

Additional checks in the report

Separate representation checks favored MTAR-L with TCR. The reported probing gains were 3.10% at Top-1000 and 8.76% at Top-50. Five-nearest-neighbor purity was 17.54% for MTAR-L, compared with 15.78% for LlamaGen-L, an absolute improvement of 1.76 percentage points.

In a supplementary RAR-B comparison, the listed FID was 18.70 for RAR-B and 17.82 for the version with MTP(B). The supplementary material also adds evaluation-protocol details, semantic-dropping schedule ablations, transfer to another autoregressive baseline, MTP-head optimization and classifier-free-guidance sensitivity analyses.

The report's scope is the class-conditional ImageNet setup at 256 × 256 described above, including the stated ablation and full-scale protocols. The document is an arXiv preprint, version 1, dated 26 Aug 2026.

Paper data and sources

Original title: Efficient Training with Foresight: Multi-Token Auxiliary Supervision for Autoregressive Image Generation
Authors: Guo Niu, Xiongfei Yao, Teng Wang, Nannan Zhu
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-26
DOI: Not available
Original paper · Full text

Versions and corrections

  1. Published automatically after legal-source, freshness, evidence, and independent-verification gates passed.