A new preprint puts AI models through a wine-specific benchmark and finds a clear split: mean accuracy was 93.6% on entry-level questions but 58.7% on expert items. Across 16 configurations, overall scores ranged from 53% to 84%, with o3 leading at 83.6%.
A test built around traceable sources
The test, called OenoBench, contains 3,266 multiple-choice questions across six wine-domain pillars and four difficulty tiers. The same release was given to all 16 configurations in one end-to-end run, with scoring based on single-letter A–D answers.
At its base is a corpus of 38,104 atomic facts collected by 35 scrapers and organized across six domain pillars and three source tiers. Language models rephrased verified facts and audited questions, but were not used as the source of truth. Each fact was tied to a URL, each question anchored to a fact and every question received a nine-agent audit.
The audit changed the release substantially. It dropped 341 questions, relabelled the difficulty of 1,259 and retained 1,601 B2-flagged questions with disclosure. A later review removed 54 defects and nine borderline-review questions, leaving 3,266 questions in the final corpus.
To calibrate the audit, three WSET-certified reviewers independently rated a stratified 50-question gold sheet for each release. The study used Cohen’s κ, a measure of agreement, and downweighted signals below 0.6.
The differences were not just about model size
Comparing standard with reasoning configurations did not show a consistent advantage. Only the DeepSeek R1–V3 comparison showed a statistically distinguishable difference: R1 was 6.8 percentage points ahead, with a 95% confidence interval from 4.6 to 8.8 points. The other three reported frontier pairs were indistinguishable from zero.
The benchmark also found a strong self-preference pattern. Anthropic configurations scored 9 to 10 percentage points higher on their own generator-family questions than on other families. Google configurations scored 6 to 10 points lower on their own, while OpenAI’s self-preference was statistically zero.
The study’s reported cost-accuracy frontier—a set of options balancing score against cost—contained five configurations. Llama 3.1 8B scored 60.5% at $0.01, Gemini 2.5 Flash 75.1% at $0.12, GPT-5-mini 78.4% at $2.82, Claude Opus 4.7 81.0% at $3.35 and o3 83.6% at $11.80.
A revealing split—and a boundary
Another comparison split the release into 1,601 B2-flagged closed-book questions and 1,665 contextual questions. Mean accuracy was 32.6 percentage points higher on the closed-book slice, with reported gaps ranging from 26.6 to 39.6 points. The authors regard the contextual slice as the more discriminating test of source-grounded reasoning.
That result is a contrast between two benchmark slices, not a direct test of practical wine expertise. OenoBench measures exact answers on a fixed multiple-choice release, and the paper does not present those scores as evidence that a model will perform reliably in sommelier, winemaking, viticulture, purchasing or other real-world workflows.
The work remains a preprint, dated 20 August 2026. The released v1.2 dataset contains 3,266 questions under CC-BY-SA-4.0, while the construction code and human-review application are reported as available on GitHub under Apache 2.0.
Paper data and sources
Original title: OenoBench: A Wine-Domain Benchmark for Knowledge-Grounded Evaluation of Large Language Models
Authors: Nikita Khudov
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-20
DOI: Not available
Original paper · Full text