Ultrasound-trained models came out on top in a computational test of image-similarity measures for B-mode ultrasound, showing the highest correlations with the confidence of a downstream classifier. Their scores most closely tracked how strongly that classifier responded in the comparison, giving the study its clearest result.
This is a result about agreement between models. The test asked whether a candidate LPIPS backbone could follow a downstream model’s confidence. LPIPS is the learned image-distance measure at the centre of that comparison.
The preprint examines whether ultrasound foundation models are more appropriate LPIPS backbones for B-mode ultrasound than models trained on natural images or broader medical imagery. The central issue is whether the data used to train a model affects how useful it is for judging differences between ultrasound images.
Putting the training data to the test
The comparison used L2 and SSIM baselines alongside candidate LPIPS backbones grouped by pre-training domain. It included models trained on natural images, radiology or general medical images, and ultrasound. CLIP, BiomedCLIP and Ultrasound-CLIP were added to examine the possible effects of both training data and training strategy.
For the supervised evaluation, the researchers used EchoGains to induce ultrasound-based augmentations on echocardiogram B-mode images. They then compared the image-distance measures with the confidence of downstream models. The setup allowed the study to ask whether a metric ranked altered ultrasound images in a way that matched task-specific model behaviour.
The classification comparison put ultrasound models at the top of the reported qualitative ranking. CNN models were also highly correlated with classifier confidence. MedSAM and USFM were among the least correlated, and ImageNet ViT outperformed them in that comparison.
That ordering is important because it separates the broad result from any claim that every ultrasound model performs the same way. The reported signal is specifically the correlation between a candidate metric and downstream classifier confidence, with ultrasound models showing the strongest relationship in the comparison.
The same question in image reconstruction
The study also tested the candidate metrics during reconstruction with an implicit neural representation, or INR, a model used to rebuild ultrasound volumes. This part used 16 follicle volumes and 75 prostate volumes, combining an L2 loss with one additional candidate LPIPS metric.
That design kept a common L2 component in place while changing the additional image-similarity term. It therefore extended the same backbone question into a reconstruction setting: whether a model trained on ultrasound offers a useful way to judge the visual differences that emerge as a volume is rebuilt.
The reconstruction datasets covered two different kinds of ultrasound volume, with the analysis reporting 16 follicle volumes and 75 prostate volumes. Those figures define the scope of this part of the comparison and show that the reconstruction test was built around specific ultrasound datasets rather than a general image benchmark.
A result about choosing the measuring tool
Taken together, the preprint offers a focused message for researchers designing ultrasound image-quality measures: the training domain of an LPIPS backbone matters enough to test directly. In this comparison, ultrasound models produced the strongest link to downstream classifier confidence, while the reconstruction arm examined the same family of choices on follicle and prostate volumes.
The finding is best read as a guide to metric selection within the study’s tested settings. It does not turn model-to-model correlation into a standalone definition of image quality; it shows which candidate backbones aligned most closely with the downstream classifier in this B-mode ultrasound comparison.
Paper data and sources
Original title: UltraPIPS: Improving model perception in B-mode ultrasound with foundation models
Authors: Tal Grutman, Tali Ilovitsh
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-26
DOI: Not available
Original paper · Full text