A preprint comparing tuberculosis (TB) cough classifiers across three public datasets found that external performance was generally weaker than performance measured within the source dataset. The study used ROC-AUC, a score that describes how well a model separates participants with TB from those without it.
A Zambia-trained deep-learning model reached a ROC-AUC of 0.755 ± 0.056 within Zambia. It scored 0.717 ± 0.124 on a Zambia audio-recorder subset and 0.741 ± 0.035 on TBscreen forced cough, but 0.632 ± 0.016 on TBscreen passive cough and 0.581 ± 0.015 on CODA. The ± figures represent variation across 10 outer-fold estimates.
The gap appeared across model types
Classical models showed a similar result. Their best within-dataset ROC-AUCs were 0.700 ± 0.053 on Zambia, 0.711 ± 0.099 on TBscreen passive cough and 0.631 ± 0.027 on CODA, while external performance was consistently weaker. Models trained on TBscreen passive cough or CODA had external ROC-AUCs below 0.6.
Researchers compared classical machine-learning and deep-learning cough classifiers across CODA, TBscreen and Zambia CIDRZ, adding analyses of country-level disease rates, recording devices and a clinical-variable baseline. Within each source dataset, nested subject-level validation used five outer folds and four inner folds, repeated twice to produce 10 outer-fold estimates. Cough-level probabilities were averaged into one prediction per subject, and external datasets were evaluated directly. Training used minority-class oversampling and upsampling across acquisition-defined subgroups.
The recordings reflected their source
Diagnostics of the representations built from the recordings showed stronger organization by dataset, recording device and location than by TB label. Within a dataset, TB-positive and TB-negative samples were often similar.
In a CODA analysis, average predicted TB probability was associated with country-level disease rate in both TB-negative and TB-positive samples. In the challenge-replication setting, the regression slopes were 1.20 for negative samples and 1.30 for positive samples, with R-squared values, which measure how closely the values fit that relationship, of 0.95 and 0.89. Under conservative preprocessing and balancing, both slopes were 0.33, while R-squared was 0.75 for negative samples and 0.84 for positive samples. The conservative setting weakened the relationship but did not remove it.
Device testing in Zambia showed that cross-device evaluation was generally lower than within-device testing. In one audio-recorder holdout, single-phone training produced ROC-AUCs of 0.675, 0.655 and 0.725, while training on all three phones produced 0.732. The individual training sets covered 632, 628 and 642 unique subjects, compared with 650 in the combined set.
A clinical baseline transferred more steadily
A logistic-regression baseline built from clinical variables transferred more consistently than the acoustic pipelines. Trained on Zambia, it reached 0.767 ± 0.025 within Zambia, then 0.655 ± 0.004 on TBscreen and 0.673 ± 0.004 on CODA. Its external performance was not uniformly high, but was more consistent than the acoustic models' results.
A warning about validation
The authors interpret the pattern as suggesting that acquisition-specific differences may encourage shortcut learning—a model picking up clues linked to a device, location or recording protocol rather than to TB. The study presents that as a plausible interpretation, not a definitive explanation, because the observed differences are tied to the datasets and settings being compared.
Specialized domain-generalization methods showed no consistent improvement on external datasets in these experiments. With only three public datasets in the comparison, strong within-dataset performance should not be treated as evidence of clinical readiness. Broader external validation across devices, sites, countries and participant groups is needed before wider transportability can be judged.
Paper data and sources
Original title: Why ML-based cough models do not generalize: a systematic cross-dataset evaluation for tuberculosis screening
Authors: Wensi Zhang, Tomas Teijeiro, Jérôme Thevenot, David Atienza
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-26
DOI: Not available
Original paper · Full text