Prompt framing was linked to the largest gap in harmful responses from 16 multimodal AI models tested in a new benchmark: the difference between the most and least vulnerable framing exceeded 40 percentage points. Task-relevant images also tracked with higher attack-success rates than no-image inputs, especially when they depicted authorization documents, identity credentials or task scenarios.
The findings come from an arXiv preprint, version 1 dated 26 Aug 2026. Its question was how harmful intent, prompt framing, visual semantics and instruction carrier relate to multimodal jailbreak vulnerability and safety behavior.
A controlled test of several moving parts
To examine those relationships, the benchmark kept the underlying harmful intent fixed and systematically changed the surrounding context in matched multimodal comparisons. The design used a taxonomy of 272 harmful intents, grouped into nine major harm domains and 18 harm scenarios. It then combined the intents with six prompt templates, five visual semantic conditions and two instruction-carrier modes.
The evaluation included eight open-weight and eight proprietary-access models. Across all 16 systems, the full combination generated 16,320 instances per model and 261,120 responses overall. The main outcome was Attack Success Rate, or ASR, defined as the proportion of responses receiving a harmfulness score of at least 4 on a 1-to-5 scale. Conditional Attack Success Rate, or CASR, measured robustness after mismatch cases were excluded, and GPT-5 was the primary judge.
The biggest swing came from wording
Average ASR ranged from 2.17% for GPT-5 to 78.38% for GLM-4.6V. That is a wide gap in measured vulnerability under the benchmark, but it should not be read as proof that one factor causes the difference across models.
Vulnerability also varied by the kind of harmful intent being tested. Cyber abuse, economic harm, privacy-related behaviors and deception-related tasks generally recorded higher ASR, while physical harm and sensitive content recorded lower ASR. The pattern means a model's response could look different across the benchmark's harm domains, another reason to treat the figures as associations within a test design rather than a universal ranking of safety.
Prompt wording was the strongest source of variation reported by the study. Story, structured and academic framings generally had higher ASR; system-style and safety-paradox framings had lower vulnerability. The distance between the most and least vulnerable framing exceeded 40 percentage points.
Images carried different signals
Images with different meanings did not behave alike. Compared with no-image inputs, the reported ASR difference was +12.96% for authorization documents, +10.47% for identity credentials and +10.10% for task scenarios. Non-semantic controls changed ASR by less than 2.2%. In these matched comparisons, the result associates vulnerability more with the semantic content of the visual context than with simply adding an image, but it does not show that rendered instructions universally increase jailbreak susceptibility.
The instruction carrier was associated with another pooled difference. With the instruction in TEXT, ASR was 50.70%, compared with 40.59% for OCR. CASR was 50.89% for TEXT and 42.34% for OCR. Mismatch rates went the other way: 0.37% for TEXT versus 4.14% for OCR, indicating that OCR more often failed the input-understanding check used before conditional safety comparisons.
Different models, different weak points
Those pooled results did not describe every model equally. Some systems had prompt-sensitivity gaps above 80 percentage points, some were more sensitive to carrier changes, and visual-semantic ranges were generally smaller but positive across models. The benchmark therefore points to distinct factor-sensitivity profiles rather than one shared weakness that appears in the same form everywhere.
To explore what might accompany the visual effect, the authors ran internal diagnostics on gemma3-12b. Authority-document and danger contexts became increasingly separated in higher-layer representations, with the sharpest rise at layer 42. Held-out authority-document contexts were displaced along the authority-minus-danger direction and showed lower attention on harm-related and visual tokens at the sensitive layer. The paper treats these as vulnerability-associated patterns, not established causal mechanisms.
A smaller audit tracked the full run
The benchmark also tested whether a smaller configuration could preserve the reported signal. It used 1,500 stratified instances while keeping the full distribution of harm domains, prompt framings, visual semantic conditions and carrier modes. Compared with the full setting, mean ASR differed by -0.06%, with a 95% confidence interval from -0.45% to 0.30%; 30 of 32 cells were within 2%, Spearman rho was 0.997 and Lin's concordance correlation was 0.999.
Automated judging was checked against human experts. On the study's QWK agreement measure, GPT-5 matched human harmfulness ratings at 0.97. A lightweight judge scored 0.95 against GPT-5. Mismatch-detection accuracy was 96.3% for GPT-5 against humans and 99.4% for the lightweight judge against GPT-5; the corresponding macro-F1 scores were 0.85 and 0.94.
The authors state that MMJailBench is released for research purposes with documentation and evaluation tools. Its value, as presented in the paper, is a repeatable way to separate prompt framing, visual meaning, instruction carrier and harm domain in matched audits. The evidence covers the 16 evaluated systems and the reported configurations, while the internal analysis centers on gemma3-12b. Whether the same factor sensitivities and internal patterns generalize beyond those settings remains open.
Paper data and sources
Original title: MMJailBench: A Factorized Benchmark for Disentangling Multimodal Jailbreak Vulnerabilities
Authors: Tianshi Wang, Jingsong Wang, Yafei Huang et al.
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-26
DOI: Not available
Original paper · Full text