An artificial-intelligence model that links CT scans with radiology-report observations recorded the highest mean AUROC among five 2D and 3D vision-language baselines across four finding cohorts. Mean AUROC, a score for discrimination, was 0.749 on CT-RATE, 0.689 on chest PMBB, 0.683 on abdominal PMBB and 0.675 on RSNA-2023.
The document is an arXiv version 1 preprint dated 26 August 2026. It describes a retrospective computational analysis of existing data, with no prospective recruitment or intervention. ACT was trained on 38,317 patients—20,000 CT-RATE chest patients and 18,317 Merlin abdominal patients—and evaluated on 29,431 scans from 25,183 held-out patients.
How the model builds its signal
ACT encoded axial slices with DINOv2, combined them with a lightweight Transformer and aligned the resulting volume representation with the paired report through a contrastive objective. Its observation bank contained 376,194 distinct radiological observations.
The observation space also showed a measurable link to a radiology term hierarchy: distances between observations tracked their separation in RadLex, with Spearman ρ = 0.565. But the test for a further increase between observations four and five steps apart was not supported, with p = 0.059.
For CT pulmonary angiography (CTPA) phenotyping—using scans to predict coded health conditions—the INSPECT test set retained 221 encounter-derived phenotypes with at least 50 positive test scans each and evaluated 2,612 scans from 2,223 patients. Those phecodes were encounter-derived weak labels, not adjudicated CTPA findings.
ACT’s concept-anchored representation reached a macro AUROC of 0.651 versus 0.572 for CT-CLIP under zero-shot scoring, which uses the representation without a newly fitted task-specific classifier. With linear probes trained on frozen representations, the scores were 0.709 for ACT and 0.662 for CT-CLIP; the paired differences were 0.079 and 0.047, respectively.
The audit exposed repeated signals
The audit ranked report-derived observations against each phenotype probe. Only 97 distinct strings occupied the 221 rank-1 positions, and an aortic/coronary calcification phrase ranked first for 20 phenotypes, including osteoporosis NOS, urinary tract infection and major depressive disorder.
The repeated ranking is a global semantic pattern, not a patient-level explanation. Because the audit was observational and global, it did not show that a confounder was used or removed in any individual scan.
Restriction kept average scores close
Researchers then used clinician-defined rules to restrict non-target observation directions and compared the resulting probes with full-bank ACT on held-out data. Across 86 eligible phenotypes, rule-restricted probes had mean held-out AUROC of 0.751, compared with 0.741 for full-bank ACT; scores were higher for 55 phenotypes and lower for 31.
The restriction comparison was descriptive: the confidence intervals overlapped, and no hypothesis tests were conducted for it. The result therefore indicates similar average discrimination in this analysis, rather than establishing that the rules improve clinical utility.
Why the result remains preliminary
These are model and representation results, not evidence of a validated clinical tool. Because the study used retrospective existing data, its audit did not establish patient-level shortcut use, prospective performance or benefit to patients. The ACT–CT-CLIP comparison also did not isolate concept anchoring as the sole reason for the performance difference.
Paper data and sources
Original title: Auditable CT Phenotyping Through Report-derived Radiological Observations
Authors: Riga Wu, Walter Witschey, Yicheng Li et al.
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-26
DOI: Not available
Original paper · Full text