A preprint reports that an AI system trained with MRI alongside partially paired tabular and PET data, plus handwriting data with no ADNI subject overlap, produced stronger held-out results for separating Alzheimer’s disease (AD) from cognitively normal (CN) participants than an MRI-only model. In the full MRI, tabular, PET and handwriting setup, the area under the receiver operating characteristic curve (AUC), a score that summarizes how well the model separates the two groups, was 0.868 with a standard deviation of 0.020. That was a 7.9-percentage-point gain over MRI-only; the adjusted p value was 0.011, and CN recall was 0.863.
The system, called PANDA in the manuscript, is built around two training stages. Auxiliary encoders first estimate a class prototype, a learned reference representation for each class. Those prototypes are then frozen, and the primary encoder is trained to align with them for all subjects, including records that lack the auxiliary modality. When the system is used, it needs only the primary modality, such as MRI, rather than all of the data used to train it.
A design built for incomplete records
The ADNI binary cohort contained 1,021 subjects: 297 with AD and 724 CN. Of these, 844 subjects, including 249 AD and 595 CN, were assigned to training and validation, while 177, including 48 AD and 129 CN, formed the held-out test set. The experiments used five-fold stratified cross-validation, three fixed random seeds and one held-out test evaluation per seed, with results reported as means plus or minus standard deviations.
Auxiliary records were unevenly available. Tabular scores were paired for about 45% of the 844 training subjects, or 378 people. FDG-PET was paired at about 19%, or roughly 158 people per training fold. The handwriting anchor had a pairing rate of zero because its participants did not overlap with ADNI.
The scanner gap narrowed, but remained
The clearest practical difference appeared when results were split by scanner field strength. For the full model, AUC was 0.783 on 1.5-tesla scans and 0.894 on 3-tesla scans. The corresponding false-positive percentages were 27.0% and 8.3%; the MRI-only comparison had a 52.2% false-positive percentage at 1.5 tesla. The extra training signals were therefore associated with fewer false alarms at 1.5 tesla, although the model did not remove the performance gap between scanner groups.
The result was also seen with a different trainability choice. On the 177-subject evaluation protocol, a fully trainable Conv5-FC3 MRI-only baseline had AUC 0.881 plus or minus 0.009, compared with 0.893 plus or minus 0.003 for the full PANDA version.
Pairing needs varied by modality
Pairing-rate tests found that the system did not always require complete matching. With tabular data alone, AUC was 0.790 at full pairing and 0.785 at 50%, but fell to 0.758 at 25%, equivalent to about 95 paired subjects. In the joint tabular-plus-PET setup, AUC stayed between 0.849 and 0.867 as pairing fell from 100% to 5%. The 5% and 10% settings matched full pairing within the reported seed noise, indicating that pairing behaved differently across auxiliary combinations.
MCI ordering was a tougher test
The preprint also tested whether the model’s output followed an ordered pattern in held-out MCI, or mild cognitive impairment, cases rather than only separating two classes. The check combined the 177 ADNI test subjects with 147 MCI subjects. MRI-only pAD had a Kruskal-Wallis p value of 0.0023 but failed the full median ordering from CN to AD. Full PANDA and its severity-head variant passed that ordering; their reported p values were 0.060 for pAD and 0.088 for the severity score, both marked nonsignificant.
The optional severity extension adds a masked-Huber severity head using complete CN/AD graded composites. It drew on 265 subjects for sorth and 378 for sdiag, while MCI subjects supplied no training gradients.
Survival results were directional, not conclusive
The cross-domain test was less decisive. In TCGA-Lung, the binary two-year overall-survival analysis used 594 subjects, including 119 in the test set; the Cox proportional-hazards analysis used 853, including 171 test subjects. With RNA paired for 100% of the relevant cases, PANDA reached an AUC of 0.628 plus or minus 0.007 and a survival-ranking C-index of 0.550 plus or minus 0.016. These represented gains of 3.5 AUC points and 9.0 C-index points over WSI-only. The 95% confidence intervals were 0.521 to 0.733 for AUC and 0.482 to 0.616 for C-index. Reported p values were 0.146 and 0.059, so neither result reached the conventional 0.05 threshold.
One part of the mechanism remains uncertain
A separate control made the method’s mechanism harder to pin down. The geometry-transfer variant changed the intended representation geometry: in the real arm, the cosine score moved from -0.597 (SD 0.020) to -0.769 (SD 0.019), while geometry loss fell from 0.412 (SD 0.035) to 0.089 (SD 0.012). Yet a four-arm ablation could not distinguish that arm from a gradient-severed no-op on downstream AUC. Pathway-specific gradient norms were not logged, leaving shared training exposure as a possible but unconfirmed explanation.
The result remains a model comparison
Taken together, the reported pattern is clearest in the held-out ADNI comparison and scanner breakdown. The TCGA figures were directionally better than WSI-only but did not reach the conventional threshold, and the geometry control did not isolate a unique downstream gain. These are computational model comparisons under the reported cohorts and splits, so the preprint does not establish clinical effectiveness or performance on unseen cohorts.
Paper data and sources
Original title: PANDA - Prototype-Anchored Alignment for Partially Unpaired Multimodal Learning, with Applications to Alzheimers MRI and TCGA Pathology
Authors: Sheethal Bhat, Mahfuzur Rahman Chowdhury, Paula Andrea Perez-Toro et al.
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-26
DOI: Not available
Original paper · Full text