AI systems built to spot made-up details in text generated from images showed only modest performance across four languages, according to a shared-task analysis. The best systems averaged 0.58 on character-level correlation, 0.46 on label-conditioned correlation and 0.51 on IoU, the measure of overlap between predicted and reference error spans. Those averages were reported as 30 to 40 points above the baselines. Even so, most mean scores were below 0.4 and no system exceeded 0.6, while Chinese Corr was slightly above 0.4 and English was described as the most challenging track.
The test looked for exact errors
The task required systems to localize hallucinated character spans—the exact stretches of text containing an error—and assign them to a five-class taxonomy: Invention, Mischaracterization, OCR Problem, Miscounting and Other.
The benchmark, called SHEEP, contains 20,000 samples in Chinese, English, French and Italian. It combines outputs from five large vision-language models with 1,600 human-written items. The dataset retained only non-synthetic images and used machine-translated prompts for French, Italian and Chinese.
The training phase provided about 15.2K annotated samples. The evaluation phase used a hidden test set of about 4.8K samples, with roughly 1.2K for each language and the labels withheld. The task received 623 submissions from 27 teams: 208 in English, 141 in Italian, and 137 each in French and Chinese.
What counted as a good detector
To judge the submissions, the benchmark used three primary measures. Corr compares reference and predicted hallucination probabilities at the character level; Corrlbl makes that correlation label-aware; and IoU measures the overlap between predicted and reference hallucinated spans. The measures capture confidence agreement, whether the error type is taken into account, and how accurately the error's location is marked.
The top position depended on the language. TÜRKSAT led English with Corr 0.55 and IoU 0.48. vroom-vroom led French with Corr 0.58 and IoU 0.52, Italian with 0.56 and 0.48, and Chinese with 0.61 and 0.53.
The leaderboard was not fixed
Those leads were less secure under statistical resampling. The analysis used bootstrap resampling to estimate rank intervals; TÜRKSAT's English result had a mean rank of 2.9, but a 95% rank interval from first to 10th. Across the language tracks, intervals for systems in the top five stretched as wide as 10 positions in Chinese, 14 in Italian and French, and 15 in English. The figures make small gaps on the leaderboard difficult to treat as definitive.
The way examples were sampled also influenced the picture. Human-written and Silver-label partitions produced similar broad rankings, with rank correlations ranging from 0.861 to 0.973, while Random sampling generally caused more reordering. Friedman tests found strategy-dependent raw scores for every metric in English and Chinese and for Corrlbl in Italian and French. They found no such Corr or IoU score difference in Italian and French.
Reading the result in context
Empty cases added another complication. Empty annotations made up 15.5% to 25.5% of test instances, while systems produced empty predictions for 30.7% to 46.0%. The evaluation did not explicitly assess whether those abstention decisions were appropriate.
The data are reported as released under a CC-BY-NC license. The document is an arXiv version-1 preprint dated 26 Aug 2026, and funding is not reported in the supplied text or metadata. Taken together with the wide rank intervals and strategy effects, the findings are best read as a benchmark comparison whose ordering needs care.
Paper data and sources
Original title: Overview of SHROOM-Visions 2026: A Shared Task on Hallucination Detection in Large Vision-Language Models
Authors: Raúl Vázquez, Aman Sinha, Chuyuan Li et al.
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-26
DOI: Not available
Original paper · Full text