Six vision-language models showed more agreement when judging the affective qualities of 3D objects—ideas such as whether a form feels luxurious or cheap—than when using unrelated adjective pairs. But their agreement was still weaker than for straightforward geometric descriptions, and it varied substantially from one object category to another.
The preprint proposes using agreement between models as a first-stage check before exposing semantic controls in generative design interfaces. It does not show that the models’ rankings match people’s own impressions of shape, because the audit was not validated against human Kansei judgments.
An audit of model rankings
The researchers tested 4,950 unique ShapeNetCore objects from 10 product categories, including chairs, tables, lamps, sofas, cabinets, bookshelves, bottles, jars, clocks and cars. Objects were randomly sampled with a fixed seed, up to 500 per category; eight categories reached that cap, while the sample contained 498 bottles and 452 bookshelves.
Each object was rendered from eight fixed viewpoints as a uniform matte-grey form, and the views were combined into a single representation. Six pretrained vision-language encoders processed the same objects using publicly released weights, with no fine-tuning.
For each adjective pair, the study created a direction between the two text meanings and projected each object’s representation onto it, producing a continuous score. The models’ agreement was then summarized with Spearman rank correlations, which compare whether they put the same objects in a similar order.
Affective rankings landed between controls and noise
Mean agreement was 0.441 for geometric adjective pairs, 0.364 for affective axes and 0.135 for irrelevant axes. In other words, models aligned most on geometric properties, reached an intermediate level of agreement on affective descriptions, and agreed least when the adjective pairs were unrelated to the objects.
Affective axes showed greater agreement than irrelevant axes, with a common-language effect size of 0.906 and a Kolmogorov–Smirnov statistic of 0.70. Geometric controls exceeded the null more strongly, with corresponding values of 0.949 and 0.81.
The pattern was not uniform across categories. Shared-core affective convergence ranged from 0.21 for bookshelves to 0.51 for jars. Full-vocabulary means ranged from 0.26 to 0.53, and the full-vocabulary and shared-pair results were correlated at 0.71.
The researchers also tested whether differences in view reliability might explain the category results. Split-half reliability averaged 0.76, ranging from 0.61 for bookshelves to 0.89 for jars. Its raw correlation with convergence was 0.47, but fell to 0.10 after correction; category differences in convergence still spanned 0.32 to 0.58.
Why the agreement may matter for design tools
Agreement was positively associated with how closely an adjective direction matched the main ways shapes varied within a category. The correlations were 0.78 for geometric axes, 0.72 for affective axes and 0.24 for irrelevant axes, with p < 0.001 throughout. The result is an association, however, not evidence that this alignment causes models to agree.
The audit also found a degree of category specificity. Affective axes generally agreed more on their intended category than on foreign categories, with mean convergence of 0.35 versus 0.28, a common-language effect size of 0.64 and p = 0.009. The pattern reversed for bottles, clocks and tables.
Agreement among five encoders was associated with the agreement of the sixth for every encoder, with an overall correlation of 0.67. This still does not answer whether the rankings correspond to human judgments.
The proposed interface would use 0.14 as a baseline from unrelated adjective pairs. A control would be exposed when its lower bootstrap bound exceeded that floor, marked provisional when its interval overlapped the null, and withheld when its estimate was at or below it. In the paper’s example, a luxurious–cheap control reached 0.47 for jars but was withheld at 0.10 for bookshelves and 0.14 for cabinets.
A consistency check, not a measure of human taste
The core limitation is direct: the audit was not validated against human Kansei judgments. Agreement between models therefore cannot be read as evidence that the scores capture how people experience shape.
The interface itself was not evaluated with users. The findings apply to the tested models, object categories, renderings and adjective vocabulary, leaving broader questions for further study.
The manuscript is identified as arXiv:2608.25876v1, dated 26 August 2026, and is presented as a preprint. The work was funded by the European Union under Horizon Europe, Grant Agreement No. 101226927.
Paper data and sources
Original title: Do Vision-Language Models Agree on the Affective Qualities of Shape? A Cross-Model Audit for Generative Design Interfaces
Authors: Luca Bux, Thiago Rios, Ingo Scholtes, Stefan Menzel
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-26
DOI: Not available
Original paper · Full text