Preprint

AI Models Put Justice First in Rare-Disease Ethics Test

An arXiv preprint found that equality-based choices made up the largest share of Justice selections, while decision-maker framing was strongly associated with the values chosen.

Justice ranked as the leading ethical value in all 11 language models tested on rare-disease dilemmas, according to an arXiv preprint. The study’s normalized win rate, used to compare how often a value prevailed, ranged from 57% to 70% across the models. Confidence intervals for that range were not reported.

Within the models’ Justice selections, Equality, defined in the study as equal resource distribution, was the largest component. Need-based and Equity-based Justice appeared less often.

A test of forced choices

The benchmark evaluated 11 large language models on 208 distinct rare-disease vignettes derived from Orphanet and OMIM data. Six reviewers with medical or biomedical backgrounds conducted a feasibility audit that resulted in 208 validated clinical vignettes.

Each model received every vignette in a standardized forced-choice exercise. The narratives concealed the ethical labels, and each prompt asked the model to select one of two actions. Researchers then mapped the A/B response to an ethical value at the model-vignette level.

The analysis compared the competing ethical values and tested whether features of the scenarios were associated with the value selected. Its association statistics describe patterns in the model responses; they do not establish that a particular feature caused those responses.

The role assigned to the decision-maker stood out

Among the contextual features examined, decision-maker framing had the strongest association with the selected ethical value. The scenario could identify a Committee, Medical Team or Individual as the decision-maker. The reported Cramér’s V was 0.504, with p < 0.001. Cramér’s V is a measure of association between categorical features, not a measure of cause and effect.

Patient type and patient age also showed associations with the selected value, but the reported measures were smaller: Cramér’s V was 0.206 for patient type and 0.181 for patient age, with p < 0.001 for both.

The regression analysis used Committee scenarios as the reference point. Against that reference, Individual contexts were associated with 5.71 times the odds of an Autonomy selection, while Medical Team contexts were associated with 3.72 times the odds. The reported 95% confidence intervals were 3.41 to 9.56 and 2.11 to 6.59, respectively, with p < 0.001 for both. These odds ratios compare selection odds with the Committee group and are not percentages of models.

Beneficence showed a similar association. Its selection had 4.23 times the odds in Individual contexts and 3.46 times the odds in Medical Team contexts, compared with Committee contexts. The reported 95% confidence intervals were 1.57 to 11.40 and 1.25 to 9.59, with p = 0.004 and p = 0.017, respectively.

The pattern was not the same for Nonmaleficence. Neither Individual nor Medical Team framing was significantly associated with selecting it relative to Committee framing. The reported odds ratios were 1.26 and 1.83, with p = 0.57 and p = 0.17, respectively.

Model identity showed little association with the selected value in the screening analysis. Its Cramér’s V was 0.067, with p = 0.416, which the analysis classified as negligible and non-significant.

What the test can and cannot settle

The benchmark’s construction limits how broadly the findings can be read. The vignettes were initially generated with GPT-4.1, the initial clinician and bioethicist critiques used during development were simulated, and the benchmark contained uneven distributions of value-pair conflicts and decision-maker contexts.

Justice-specific regression estimates were considered unreliable because the decision-maker groups were heavily unbalanced: 324 Committee contexts, 27 Medical Team contexts and six Individual contexts.

The manuscript is an arXiv preprint, version 1, dated 25 Aug 2026. The authors state that the final dataset of 208 vignette JSON files, along with the associated curation and generation code, will be released upon publication acceptance.

Paper data and sources

Original title: Rare Diseases, Common Dilemmas: LLMs Prioritize Equal Resource Distribution over Patient Benefit in Decision-Making
Authors: Minda Zhao, Xu Han, Rishabh Goel et al.
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-25
DOI: Not available
Original paper · Full text

Versions and corrections

  1. Published automatically after legal-source, freshness, evidence, and independent-verification gates passed.