Preprint

Preprint: Medical AI models lose accuracy in lower-resource languages

A nine-language benchmark found the steepest gaps in Swahili and Zulu, while proprietary systems were more stable than open-source and medically specialized models.

A preprint benchmark of medical large language models found that average accuracy on multiple-choice and natural-language-inference tasks was 15.4 percentage points lower in Swahili and 28.9 points lower in Zulu than in English. The five proprietary models ranked in the top five overall and had resource gaps of 1.9 to 5.5 points, while the three medically specialized models had gaps of 13.6 to 20.8 points.

Known as HealMed, the benchmark contains 1,000 aligned examples in each of nine languages, drawn from nine datasets and covering MCQA, NLI and open-ended QA. It was developed over two years by 23 physicians and medical experts based across nine countries and regions.

The biggest losses came in lower-resource languages

The study defined a resource gap as the difference between mean accuracy in higher-resource and lower-resource language groups. It first macro-averaged accuracy across datasets within each language and then across the nine languages.

For MCQA and NLI, the proprietary models had lower-resource mean accuracies of 79.0% to 84.0%. The highest lower-resource mean among open-source and medically specialized models was 60.0%.

Accuracy was lower than English in all eight translated languages. Reductions were 2.1 to 3.8 points in German, Spanish and Portuguese; 6.1 in Japanese, 6.5 in Chinese and 5.9 in Thai; 15.4 in Swahili and 28.9 in Zulu.

The pattern continued in open-ended answers

For open-ended QA, an LLM judge scored answers for completeness, alignment with a reference answer, clinical consensus, clinical appropriateness and safety. Each dimension was rated from 1 to 5, and the five ratings were averaged into an overall score.

The proprietary models were highest-scoring and most stable in this part of the test, while all eight non-proprietary models declined in lower-resource languages. GPT-5.4 scored 4.31 in higher-resource languages and 4.35 in lower-resource languages; Gemini-3-Flash scored 3.93 in both groups.

Among general-purpose open-source models, observed declines included 1.35 and 1.20 points. The medically specialized models showed gaps ranging from 0.50 to 1.38 points.

Translation was part of the measurement

The researchers compared machine-translated versions with versions reviewed by bilingual medical experts. Each target-language instance went through a two-stage review, with experts rating accuracy, fluency and completeness on a five-point scale.

After English adjustment, the mean absolute accuracy shifts were 4.9 points in Thai, 3.8 in Swahili and 5.8 in Zulu. Thai's largest shift was 6.9 points on MMLU-Pro; selected NLI shifts were +3.4 points in Chinese MedNLI, +3.3 in Swahili MedNLI and +3.2 in Zulu BioNLI.

For open-ended QA, machine-translated versions generally received higher judge scores than expert-reviewed versions. The average absolute shifts ranged from 0.06 to 0.12 points, and expert-reviewed scores were lower in 18 of 24 language-dataset combinations; in Chinese LiveQA, expert review corresponded to a 0.15-point lower score.

The result is not a blanket verdict on translation: the direction and size of differences varied by language and dataset. The authors also noted that the automated judge may have been aligned with literal machine-translated wording.

Across eight target languages, reported mean translation ratings were 4.72 for accuracy, 4.71 for fluency and 4.88 for completeness. The mean word-level revision rate was 4.8%, with language rates from 1.0% to 12.1%; the authors cautioned that separate expert groups and different correction thresholds limit direct ranking of languages.

The automated scorer did not always match experts

To test the automated judge, the validation used 15 score-stratified questions per language—five from each of the three QA datasets—and three models, producing 45 paired evaluations per language. The validation covered selected languages and questions rather than the full benchmark.

Within those subsets, LLM scores fell within one point of expert scores for 95.6% of Thai evaluations, 88.9% of Chinese evaluations and 64.4% of Japanese evaluations. The corresponding concordance coefficients were 0.77, 0.56 and 0.11, and the paper says automated and human evaluation were not interchangeable.

A benchmark, not a clinical verdict

These are benchmark results based on aligned translated items and model-generated answers. They describe performance under the included languages, tasks, models, prompts and scoring procedures; they do not establish patient-level clinical safety, effectiveness or usefulness in real-world care.

The comparisons were descriptive, and no inferential tests, confidence intervals or formal uncertainty estimates were reported.

Paper data and sources

Original title: HealMed: Multilingual Evaluation of Large Language Models in Medicine
Authors: Yingjian Chen, Fan Gao, Sherry T. Tong et al.
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-20
DOI: Not available
Original paper · Full text

Versions and corrections

  1. Published automatically after legal-source, freshness, evidence, and independent-verification gates passed.