An AI image-generation system scored higher than its main comparator when it was tested on combinations of imaging factors that had not appeared together during training, according to an arXiv preprint. X-MULTI recorded an average I-FAA score of 0.53 on those novel combinations, compared with 0.47 for MULTI. The 0.06 gap was reported as an 11% relative difference. I-FAA is the study's measure of factor adherence. It uses class-balanced training and factor-specific augmentations to reduce shortcut learning from correlated imaging factors. In plain terms, the measure is intended to make it harder for a classifier to use one linked feature as a stand-in for another. The paper is an arXiv preprint, version 1, dated 25 August 2026; no journal publication is reported.
A test of unseen combinations
The study asks whether imaging factors can be controlled independently enough in text-to-image generation to make valid combinations not seen in training. It uses DF-RICO, a benchmark spanning 15 autonomous-driving and surveillance datasets, with annotations for lens, source or domain, viewpoint and sensor modality. MULTI was the primary comparator. The evaluation also included SDXL Zeroshot, DreamBooth and Inspiration Tree, as well as comparisons in which ControlNets were used or not used.
What X-MULTI adds
X-MULTI augments MULTI with zero-shot supervision from a vision-language model, or VLM, an image-and-language model. That supervision is part of the system's approach to factor disentanglement for novel factor combinations. The implementation used SDXL with 15 embedding vectors per factor and trained for 10 epochs with a total batch size of 4. The VLM branch began at epoch 3 with supervision strength 10^-6. Each mini-batch contained three real samples and one synthetic sample. Evaluation used I-FAA, FID, CLIP Score, Inception Score and Diversity Score alongside comparisons with the named systems.
Where the reported gains appeared
On factor-level comparisons within the novel-combination test, the reported relative improvement was 21% for lens and 18% for viewpoint. Sensor and domain each showed a 0.01 difference. Across evaluations with and without ControlNets, the paper reported X-MULTI as the best overall method on I-FAA, factor CLIP alignment and diversity. The other image-generation metrics were mixed across methods, so the reported lead did not appear on every measure.
The supervision was uneven
The additional supervision had uneven performance across the factor categories. Its average classification accuracy was approximately 0.78, but predictions for the rgb-thermal sensor and for viewpoints other than front were unreliable. Supervision was disabled for rgb-thermal and for all viewpoints except front. A separate diagnostic on ground-truth images reported average accuracy of 0.78 for Qwen2-VL-7B-Instruct, versus 0.51 for LLaVA-1.6.
Reported variation across prompts
Accuracy varied across the prompt formats used in the VLM diagnostic. Detailed factor-specific prompts reached 0.78, compared with 0.70 for simple prompts and 0.63 for one unified prompt. The supervision-strength comparison reported I-FAA of 0.53 with moderate 10^-6 and weak 10^-8 supervision, while strong 10^-4 supervision reached 0.47.
The measure was also tested
On ground-truth DF-RICO images, I-FAA's average classifier accuracy was 0.95 versus 0.43 for FAA, with the largest gains reported for viewpoint and domain. The I-FAA design uses class-balanced training and factor-specific augmentations to reduce shortcut learning from correlated imaging factors. In supplementary tests of factor combinations that already existed in the data, X-MULTI had the best I-FAA while remaining comparable on other image-generation metrics.
Residual links remain
The metric analysis found that I-FAA reduced most cross-factor correlations, but residual correlations remained for factor pairs that were strongly coupled or sparsely represented in DF-RICO. Taken together, the paper reports higher alignment for X-MULTI on the tested unseen combinations while also reporting unreliable categories and residual cross-factor correlations.
Paper data and sources
Original title: X-MULTI: VLM-based Imaging Factor Disentanglement for Factor-Aware Image Synthesis
Authors: Sonali Godavarthy, Matthias Neuwirth-Trapp, Tim-Felix Faasch et al.
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-25
DOI: Not available
Original paper · Full text