An arXiv preprint dated 26 August 2026 reports that sampling under exact fairness accounts for a substantial part of measured race and age differences in chest-radiograph model performance, but many model-finding combinations still exceed that reference. Using AUROC, a ranking-based performance measure, the median race difference was 0.082 against a fair-model reference of 0.034; for age, the figures were 0.066 and 0.014. The reference accounted for median shares of 41% of the race difference and 22% of the age difference, while 79 race combinations and 116 age combinations exceeded it.
A benchmark for sampling variation
The method, FRAME, is a two-step audit. It first estimates the subgroup difference expected under exact fairness and then tests the remaining difference using the available feature representation.
The fair-model reference came from 2,000 draws at seed 0, with the interval set from the 2.5th to the 97.5th percentile.
The primary corpus contained 702,206 images: 650,207 chest-radiography images, 41,999 dermatology images and 10,000 retinal funduscopy images. The main chest-radiograph panel used 130 combinations of 10 frozen encoders and 13 findings on a pooled test split of 125,992 images from 41,621 patients.
The result changed with the metric
In an audit of 89 claims from nine published studies, 41 exceeded the reference. Of 53 rate differences, 40 exceeded it, with a median reference share of 25%; of 36 AUROC differences, one exceeded it, with a median share of 70%. In 22 claims, the reference was larger than the reported difference.
Representation tests offered no simple answer
On RAD-DINO, the decodability-injection condition changed linear race decodability from 0.866 to 0.942 and nonlinear decodability from 0.857 to 0.941, while the race difference stayed at 0.077. At 256 dimensions, the entanglement-injection condition had a race difference of 0.118 versus 0.077, alongside a reported 0.035 fall in overall disease AUROC.
Across 130 combinations, linear decodability, erasure cost and principal-angle overlap had Spearman correlations of 0.036, -0.142 and -0.027, respectively, with the achievable difference; none was significant, with all pFDR values at least 0.329. Cross-fitted out-of-sample R2 was negative.
Mitigation shifts were modest
Across nine mitigation methods, the median achievable race difference was 0.071, a median reduction of 0.005; the change was significant in 10 of 128 combinations. For age, the median difference was 0.057, with a reduction of 0.008, and the change was significant in 45 of 130 combinations.
Retraining under two further pretraining seeds changed the race difference by a median 0.012, compared with a median 0.005 reduction under the nine mitigation methods. In the tested configurations, seed-to-seed variation was therefore larger than the typical mitigation change.
Thresholds moved rates more than rankings
Per-group operating-point shifting, which changes decision thresholds, left the median race AUROC difference at 0.000 and reduced median sensitivity and false-positive-rate differences by 0.075 and 0.069. Platt calibration also left the median race AUROC difference at 0.000 while reducing the calibration difference by 0.077.
Across 260 combinations, calibration left the AUROC difference unchanged in 252. The five exceptions occurred in race comparisons whose smallest subgroups had 20 to 34 positive cases.
A different training signal
In 52 paired comparisons, image-text pretraining was associated with higher worst-group AUROC than self-supervised pretraining: median gains were 0.049 for race, 0.054 for age and 0.044 for sex. The median change in subgroup difference was at most 0.015.
What the preprint cannot yet explain
The preprint does not identify what creates the remaining representation-related difference. Its interventions acted on cached features rather than during pretraining, so they do not establish a training-time mechanism.
Race comparisons carry a data complication: of 68,521 pooled test radiographs assigned to the 'Other' category, 52,038 came from sites that did not record race, so race and site are partly confounded.
The authors recommend applying the fair-model reference before selecting an intervention. The sequence is intended to show whether a measured subgroup gap is largely within the range expected from sampling or leaves a remainder for separate investigation.
Paper data and sources
Original title: FRAME: separating sampling variation from representational cause in medical imaging fairness
Authors: Mahshad Lotfinia, Daniel Truhn, Andreas Maier, Soroosh Tayebi Arasteh
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-26
DOI: Not available
Original paper · Full text