An arXiv preprint reports a computer model that maps myocardial scar in LGE-CMR images with an overall Dice score of 0.677 ± 0.24 and a scar-volume error of 38.88. Dice is an overlap measure: it compares the model’s scar map with the reference annotation.
For samples marked low confidence, CalcSeg reported a Dice score of 0.644 ± 0.22 and a scar-volume error of 35.03. These are segmentation benchmarks on image data; the study does not evaluate improved diagnosis, treatment or patient outcomes.
How CalcSeg approaches uncertainty
CalcSeg combines latent 3D context with slice-wise self-attention and a confidence-aware curriculum. The training process expands from more confident samples to ambiguous ones, bringing harder cases into the learning sequence.
To estimate uncertainty, the model makes repeated stochastic runs and measures how much its outputs vary at the final network layer. The study calculated the estimate from the mean standard deviation of 100 Monte Carlo Dropout samples for each subject.
A benchmark across several datasets
The dataset brought together LGE-CMR data from four sites and two segmentation challenges. It included 976 patients, with between four and 18 slices per subject.
Clinical experts supplied left-ventricle and scar annotations, and the data were split into training, validation and test sets at a 7:1:2 ratio. CalcSeg was compared with five named baselines: TransUNet, AttentionUNet, UNETR, ScarNet and ScarNet with supervised curriculum learning.
The main measures were the Dice overlap score and percentage error in estimated scar volume, calculated on an independent test set and on six cases clinicians had flagged as challenging.
What the reported results show
The preprint says qualitative examples showed greater overlap with the reference scar and minimal false positives than the baseline networks. Its component tests also reported better overall segmentation when slice-wise self-attention was included; the combination of self-attention and curriculum learning had the best overall accuracy and difficult-case scar-volume error.
For the clinically challenging subgroup, the reported test-set ratio was 22:6, and uncertainty fell substantially across curriculum stages. The report does not give the size of that decrease or formal inferential uncertainty for the comparison.
The report does not define what the ± spreads in the Dice results represent, and it provides no p-values, confidence intervals or formal significance tests. The figures should therefore be read as reported benchmark metrics, not as a formal statistical demonstration of superiority across settings.
Where the evidence stops
The challenging-case subgroup consisted of six cases flagged by clinicians, so its findings may be exploratory rather than a broad measure of performance across difficult scans.
The evaluation stayed within reported image datasets, segmentation metrics, qualitative examples and model-derived uncertainty. It did not test prospective clinical deployment or patient outcomes.
Open questions include whether the findings replicate on independent external cohorts, whether the uncertainty estimates are calibrated for clinical use, and how the approach would perform with multi-sequence CMR.
Funding and disclosure
The document is an arXiv preprint dated 20 Aug 2026. The authors state that the code is released on Github and report no relevant competing interests. The work was supported by NIH 1R21EB032597, the iPRIME Student Fellowship Award and NSF CAREER 2239977.
Paper data and sources
Original title: CalcSeg: Confidence-aware 3D Latent Context Curriculum Learning For Myocardial Scar Segmentation From Single-Stack LGE-CMRs
Authors: Nivetha Jayakumar, Hannah Kim, Amit R. Patel, Miaomiao Zhang
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-20
DOI: Not available
Original paper · Full text