A preprint reports that PelviNeXt reached 92.00% accuracy when classifying normal and abnormal pelvic ultrasound images—but only after an integrity audit dramatically reduced the benchmark. The PCOSGen pool fell from 4,668 images to 225, a reported 95.2% reduction.
The audit first removed exact duplicates, then grouped near-duplicates using Hamming distances. Pairs at a distance of 14 or less were treated as near-duplicates, while a distance of 16 was judged visually distinct. The final set contained 63 normal and 162 abnormal images.
Strong scores on two types of scans
PelviNeXt was applied without task-specific changes to both pelvic ultrasound images and pelvic X-rays. Its hybrid design combines dense feature extraction, hierarchical CBAM, multi-scale fusion and talking-heads multi-head self-attention.
On the deduplicated PCOSGen set, PelviNeXt recorded 91.74% recall, 86.48% specificity, an F1 score of 0.8890 and an AUROC of 0.9051, alongside its 92.00% accuracy. The reported 95% confidence intervals were ±1.60 percentage points for accuracy and ±0.0156 for AUROC.
The paper reports that PelviNeXt had the highest result across all five PCOSGen measures it compared: accuracy, recall, specificity, F1 score and AUROC. Its reported advantage over the listed comparison models was at least 2.67 percentage points in accuracy and 0.0422 in AUROC.
On PXR150, which contained 100 fracture cases and 50 normal cases, the same architecture reached 87.33% accuracy, 89.00% recall, 87.00% specificity, an F1 score of 0.8774 and an AUROC of 0.8920. The reported 95% confidence intervals were ±2.45 percentage points for accuracy, ±5.71 for recall, ±3.92 for specificity and ±0.0288 for AUROC.
Against the strongest listed prior PXR150 result, PelviNeXt was reported to perform better on accuracy, recall, specificity and AUROC. The reported specificity was 87.00%, compared with 82.00% for the earlier result.
What changed when parts were removed
The models were trained from scratch for 30 epochs, using batches of 16 images and five-fold stratified cross-validation. Augmentation was applied only to the training folds, and the reported metrics were means with 95% confidence intervals across the folds.
The researchers also tested stripped-down versions of PelviNeXt. Each altered version had lower reported performance; the largest AUROC decrease on both tasks was seen in the version without multi-scale fusion, whose AUROC was 0.8735 for PCOS classification and 0.8576 for fracture classification. The version without hierarchical CBAM had accuracy 1.56 percentage points lower on PCOSGen and 2.66 points lower on PXR150.
Results still need outside testing
The strongest figures came from small datasets: 225 deduplicated PCOSGen images and 150 PXR150 radiographs. The authors say those sizes limit statistical power and identify evaluation across different datasets and sites as future work.
The study used cross-validation rather than a reported independent external test set. The figures therefore describe performance under the reported benchmark protocols, not how the model would perform on a separate dataset or at another site.
The near-duplicate threshold for PCOSGen was chosen by visual inspection, leaving open the question of how other thresholds would change the performance estimate. The deduplicated benchmark is publicly available, while the original PCOSGen and PXR150 datasets remain available from their respective sources.
A preprint evaluation
The document is an arXiv preprint that states it was accepted at MICCAI CAPI-WOMEN 2026. The authors declare no competing interests.
Paper data and sources
Original title: PelviNeXt: A Modality-Agnostic Hybrid Network for Pelvic Imaging in Women's Health
Authors: Siam Tahsin Bhuiyan, Rashedur Rahman, Sefatul Wasi et al.
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-20
DOI: Not available
Original paper · Full text