A benchmark of six language models found that supplying ideal factual context was associated with higher accuracy on questions about public companies, but it did not bring results to a common level across global markets. Markets where models performed more strongly without supplied context generally remained stronger with perfect context. The finding tests a central promise of retrieval-augmented generation, in which a model receives outside text while answering, but the study does not establish a causal effect on geographic equality.
The same benchmark recorded a separate risk when the supplied passage was false but coherent: models often adopted the incorrect evidence, and accuracy could fall below the no-context baseline based on parametric recall, or information already encoded in the model. Irrelevant passages showed a generally weaker but systematic degradation. These were synthetic benchmark conditions, not observations of poor search results in a production retrieval system.
A test built around company facts
The benchmark covered 2,135 unique publicly listed companies across 15 major global equity indices. Because some companies appeared in more than one index, the dataset contained 2,165 index-linked records.
Researchers converted each fact into two paired multiple-choice questions: one inductive question moving from a company to an attribute, and one deductive question reversing that direction. Each question had one correct answer and three distractors, or incorrect alternatives. The exercise generated approximately 15,000 to 17,000 questions.
The questions were tested with no context, perfect context, misleading context or distraction context. The six models were GPT-5, GPT-5 mini, GPT-5 nano, Claude Sonnet 4, LLaMA-70B and LLaMA-8B. Responses were coded as correct or incorrect, with accuracy, standard errors and confidence intervals reported; the analysis also compared context-sensitive changes with the no-context baseline.
The differences were visible before context was added
Without supplied context, accuracy varied substantially by index across the models. Performance was higher in large, English-dominant markets and lower in smaller or less-represented markets, showing that the benchmark’s geographic differences were present in the starting results.
The direction of the question also mattered. Inductive questions were reported as easier than deductive questions, with the largest gap in smaller models and a narrower gap in larger ones.
Perfect context was associated with higher observed accuracy across the models, but the gains were uneven. Stronger-baseline markets generally retained their advantage, so the market-level results did not converge to a common accuracy level.
False passages were often treated as evidence
The misleading-context condition highlighted a different failure pattern. When models were shown false but plausible evidence, they often adopted it rather than rejecting it. In some model, attribute and baseline-knowledge combinations, accuracy was lower than under parametric recall alone.
Distraction context was generally less damaging than misleading context, but the decline was still systematic. Stronger models and indices with stronger baseline performance were generally more robust. The authors also note a possible English-centric embedding bias in how distractors were selected, which may have affected the difficulty of irrelevant passages across markets.
More scale improved scores without removing the pattern
Larger models outperformed smaller models across the tested conditions and were more robust to misleading or distracting evidence. But in this six-model benchmark, scale did not remove market disparities, the tendency to rely on context, or broader contextual brittleness.
The analysis also examined whether company size explained the market pattern. Revenue quintile was a significant predictor of accuracy in all four context conditions, with p-values below 0.01. Index effects remained jointly significant after adjustment for company size, with p-values below 0.001. Only one of the 15 indices showed strictly monotonic accuracy across revenue quintiles.
What the benchmark leaves open
The evidence is limited to English multiple-choice questions about atomic company facts drawn from Wikipedia-derived fields and summaries. The misleading passages were synthetic, coverage and distractor difficulty could vary by market, and the models came from US or European developers. The findings therefore do not directly evaluate production retrieval sources, non-English questions, non-Western models, people or organizations, or long-form and multi-hop tasks.
The strongest perfect-context comparison also relied on a filtered subset of questions. Filtering excluded 61.0% of Founding Year questions, 29.6% of Headquarters questions and 86.0% of Key People questions. The study’s overall interpretation is therefore conditional: retrieval may act as a corrective, but its observed benefits were not universal across markets or context conditions.
The benchmark dataset is reported as available through Zenodo. The paper’s front matter states that it was published as a conference paper at COLM 2026, while the metadata labels the item a preprint. The acknowledgments name discussion contributors but no funding source.
Paper data and sources
Original title: When RAG Fails to Equalize: Geo-bias in Factual Question Answering over Public Companies
Authors: Abhinav Havaldar, Enrico Santus
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-26
DOI: Not available
Original paper · Full text