Preprint

Concept-Based AI Outperformed No Support, but Did Not Consistently Beat Label-Only AI

Preprint: Two experiments found higher accuracy with concept displays than no support in email and bird classification, while interactive controls showed no consistent full-sample advantage over a label alone.

Concept-based AI support was more accurate than no support in two binary classification tasks, but the study did not find a consistent full-sample advantage over a label-only system.

The results are reported in an arXiv version 1 preprint dated 26 August 2026. The paper tested concept bottleneck models, or CBMs, which make predictions through intermediate concepts that users can inspect. It asked whether interactive CBMs improve human-AI team accuracy, whether users intervene when concept detection is incorrect, and whether CBM support changes confidence or trust relative to non-interpretable AI support.

What the researchers compared

Across the two user studies, the authors report 705 participants and 6,959 observations. The comparison used four between-subjects AI-support conditions: no support, a label only, a label plus concepts without editing, and the same concept display with an option to intervene. In the bird study, the first 400 recruits were assigned completely at random, while the remaining 151 were assigned to specified conditions to ensure at least 85 eligible participants per condition. Each participant classified 10 items, and supported participants saw a model that was correct on eight and wrong on two.

Study 1 used valid and phishing emails, while Study 2 used images of Le Conte’s and Savannah sparrows. Each CBM used six bottleneck concepts, an independently trained frozen neural encoder, one binary SVM, essentially a yes-or-no classifier, for each concept and a shallow logistic-regression predictor. Before the user studies, the email CBM achieved 92.3% test accuracy and the bird CBM achieved 81.4%.

A simulation-based a priori power analysis targeted 340 participants, or 85 per support condition, for 82% power. Study 1 collected 401 participants and retained 363 after exclusions; Study 2 collected 551 and retained 342. Restricting the bird analysis to the initial 400 recruits reportedly did not alter the results. The researchers used mixed-effects models that accounted for participant and item differences and applied Bonferroni correction for multiple comparisons.

The task changed the result

In the email task, observed accuracy was 76% with no support, 81% with label-only support, 83% with non-interactive concepts and 83% with interactive concepts. Each AI-supported condition was significantly higher than no support, while the three supported conditions did not differ from one another.

When the CBM label was wrong, accuracy was 68% with non-interactive concepts and 55% with label-only support. Support condition interacted with CBM correctness, with p = .002, and the non-interactive-versus-label-only comparison was p = .047.

The bird task showed a similar advantage for concept displays over no support: accuracy was 73% without support, 79% with label-only support, 82% with non-interactive concepts and 83% with interactive concepts. The two concept conditions were significantly above no support, but label-only support did not differ from no support, and neither concept condition differed significantly from label-only support in the full sample.

On items where the CBM was correct, all AI-supported conditions were more accurate than no support: 85% for label-only, 90% for non-interactive concepts and 89% for interactive concepts, compared with 74% without support. The difference between support conditions also depended on whether the CBM was correct, with p < .001.

Controls did not guarantee a gain

The interactive email group made an average of 6.87 interventions across 10 items. Intervention was recorded on 22% of mismatch trials, compared with 9% of matching trials. Intervention trials were 82% accurate versus 83% without intervention, a difference that was not statistically significant, and 31 of 87 participants, or 36%, never interacted.

The bird interactive group intervened 9.62 times on average across 10 items. It intervened on 33% of mismatch trials and 5% of matching trials; intervention trials were 88% accurate versus 79% for non-intervention trials. These within-condition figures show an association between intervention and accuracy, not proof that the controls caused the difference.

Among bird-task participants in the interactive condition who used the controls at least once, the analysis excluded 30 of 86 people who never interacted. The remaining participants had 86% accuracy, significantly above label-only support, with p = .049; on CBM-wrong items, the figures were 65% and 60%, and on CBM-correct items they were 91% and 89%. This active-user comparison cannot isolate the effect of choosing to intervene from differences between people who did and did not use the interface.

Confidence and trust did not move together. In the email task, support did not significantly change confidence or overall trust. After participants who did not interact in the interactive condition were excluded, the capability rating was lower for interactive than label-only support, with medians of 3.5 and 4.0. In the bird task, interactive support had a median confidence rating of 4 versus 3 with no support, but overall trust did not differ. In the corresponding excluded subgroup, capability medians were 4 for interactive support and 5 for label-only support.

A result with narrow boundaries

An exploratory CUB-only analysis examined concept-detection errors and adherence to correct model predictions. As errors increased, the odds of following a correct CBM prediction fell in the non-interactive condition, with an odds ratio of 0.72, while the interactive condition showed little change, with an odds ratio of 1.12. The association does not establish that inaccurate concepts reduce trust.

The authors’ conclusion is conditional: they see potential CBM performance benefits beyond unaided humans and non-interpretable AI support when users are uncertain, concepts are objective and understandable, and users engage with interaction. But the evidence covered only binary classification, did not systematically investigate concept-detection accuracy, and leaves multi-class cognitive load, unreliable concept annotations and trust effects unresolved.

Taken together, the full-sample results support a narrower reading. Concept displays were more accurate than no support in both tasks, but supported conditions did not differ from one another in the email study, and neither concept condition differed significantly from label-only support in the full bird sample. The preprint leaves open whether the pattern will hold in multi-class tasks or when concept-detection accuracy is systematically varied.

Paper data and sources

Original title: Are Concept Bottleneck Models Effective as Decision-Support Systems?
Authors: Alessandro Bogani, Nicola Debole, Emanuele Marconato et al.
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-26
DOI: Not available
Original paper · Full text

Versions and corrections

  1. Published automatically after legal-source, freshness, evidence, and independent-verification gates passed.