Human reviewers labeled 18 images in the people and hand categories as defect-free, compared with 136 images depicting objects and scenes. These are descriptive counts from the study's image categories.
The study set out to examine how people identify defects in AI-generated images with complex compositions, including multiple entities and attributes, spatial relationships and interactions. It also asked whether those human judgments could support a small experiment in defect prediction and image repair.
A dataset built around complex compositions
Researchers assembled 651 reference images: 166 featuring people, 161 focused on hands, 161 showing objects and 163 depicting scenes. They manually refined the prompts to specify multiple entities and attributes, spatial relationships and interactions.
The 651 AI-generated images were allocated to 11 training images, 40 pilot-study images and 600 main-study images. Model selection began with 80 randomly sampled compositional prompts. Midjourney, Imagen and FLUX were selected through perceptual comparison, after which the 651 prompts were randomly divided into three model-assigned groups.
How a defect was recorded
Participants visually inspected each generated image and either marked it as having no noticeable defect or reported defects at a global level, a local level, or both. Here, global meant a broad image problem, while local referred to a flaw tied to a particular location.
Fifteen people were recruited for the pilot study and 29 for the main study. Reliability checks covered engagement, self-consistency and conformity; six participants were identified as outliers, and data from 23 participants were used for subsequent analysis.
The annotation process produced 18,906 defect annotations, including 6,626 global-level annotations and 12,280 local-level annotations. The records therefore captured both broad and location-specific judgments.
The model comparison needs context
Among the assigned image groups, Imagen had nearly 70 images labeled defect-free. Half of its images were labeled Global defect, while Midjourney had 72 images labeled Local defect.
These figures describe the labels in the study's assigned groups, not a general defect rate or a causal ranking of the systems. The three models received separate prompt groups, so the comparison was not clearly paired image by image.
A repair test built from human maps
The proof of concept built ground-truth maps from human local-defect locations. It fine-tuned TranSalNet to predict similar locations, compared its output with GPT-Image-1, and used the predicted maps to guide GPT-Image-1 repairs.
In four sampled examples, fine-tuned TranSalNet produced defect-localization patterns the authors described as more human-like than those from GPT-Image-1. The defect-guided repaired images were also described as correcting defects and showing noticeable quality improvement in those examples.
The repair and localization result is illustrative rather than a broad performance test. The proof of concept used four sampled examples, and no quantitative localization or repair metric was reported. The study therefore does not show that the approach improves image quality reliably across the full dataset.
An early release of the dataset
The document is a version 1 arXiv preprint dated 26 August 2026. It presents a manually curated dataset with human annotations and qualitative proof-of-concept experiments.
The paper states that the CO-AID database and supplementary materials are available through its GitHub repository. The dataset brings together human annotations of generated images across people, hands, objects and scenes, while broader tests across other prompts, categories and models remain open.
Paper data and sources
Original title: When Composition Doesn't Add Up: Humans Identifying Defects in AI-Generated Images
Authors: Ruoqi Hu, Chulin Zhao, Jiashuo Chang et al.
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-26
DOI: Not available
Original paper · Full text