A reconstruction model using artificial intelligence has reported better image quality and velocity accuracy than four comparator methods in a benchmark of highly accelerated 4D flow MRI, including tests at 50-fold acceleration. The study, published as an arXiv version-1 preprint dated 26 Aug 2026, asks whether FlowMoDL can recover both anatomical detail and phase-derived velocity information from undersampled scans.
FlowMoDL was judged on two parts of the scan: anatomical magnitude and phase-derived velocity. The reported endpoints were structural similarity and normalized root mean square error (nRMSE), a normalized measure of reconstruction error, for magnitude images, plus relative and angular errors for aortic velocity.
A model built around the scan data
FlowMoDL is an “unrolled” model, meaning that it repeats learned and data-consistency steps in a fixed sequence. Its learned component is a (3+1)D spatiotemporal denoiser, while its correction step uses conjugate-gradient updates based on the SENSE forward model. The network also uses dual-pathway conditioning on the acceleration level.
The training objective added penalties for velocity, relative error and angular error to align the phase-derived information. It used a curriculum that emphasized magnitude reconstruction first before increasing the weight of the phase-related terms. The configured model used eight untied cascades, 10 conjugate-gradient iterations per cascade, 64 denoiser channels and four residual blocks, with 3×3×3 spatial kernels and a length-5 temporal kernel.
What the benchmark found
The researchers tested the method on the multi-center, multi-vendor CMRx4DFlow challenge dataset. Its 138 original training cases were divided into 96 training cases, 21 validation cases and 21 unseen test cases, while preserving the distributions of centers and scanner architectures. Validation and testing used pre-generated undersampling masks spanning acceleration factors from 10× to 50×, and the learned models were trained for an equivalent budget corresponding to 50 epochs.
Averaged over the tested acceleration factors of 10×, 20×, 30×, 40× and 50×, FlowMoDL reported a normalized root mean square error of 0.0469 ± 0.0017 and a structural similarity score of 0.9409 ± 0.0026. For aortic velocity, its relative error was 0.2664 ± 0.0065 and its angular error was 24.5639 ± 0.2458. The quantitative results were averaged across five random seeds.
The paper reports that FlowMoDL strictly outperformed CG-SENSE, FlowVN, FlowMRI-Net and MoDL on all four metrics across the tested acceleration factors, with lower image and velocity errors and higher structural similarity. In the aggregate comparison, MoDL reported an nRMSE of 0.1147 ± 0.0035 and SSIM of 0.8062 ± 0.0095, compared with FlowMoDL’s 0.0469 and 0.9409.
The sharpest test came at extreme undersampling
At 50× acceleration, the researchers reported that FlowMoDL retained sharp spatial structures and smooth transitions in the reconstructed velocity signal. Two temporal-coherence measures were close to the ground-truth values: lag-one temporal autocorrelation, which compares neighboring time points, was about 0.72 for FlowMoDL versus about 0.71 for ground truth, while the fraction of temporal power above half the Nyquist frequency was about 0.07 versus about 0.09. The authors described only marginal over-smoothing in this comparison.
The study also examined how the systems behaved when training was limited to an equivalent number of gradient steps. Under that selected budget, the authors report that the flow-specific comparator networks degraded significantly, while FlowMoDL converged robustly and outperformed the competing models. The result is tied to the paper’s normalized training-budget protocol.
A benchmark result, not a clinical verdict
The findings remain limited to a computational reconstruction benchmark using the CMRx4DFlow cases and the study’s custom data split. The reported outcomes are reconstruction surrogates—image similarity, reconstruction error, velocity error and temporal-coherence measures—rather than patient outcomes, diagnostic performance or clinical decisions. No inferential tests, confidence intervals or p-values were reported, and the learned-model comparison depended on the equivalent gradient-step budget and 50-epoch setup.
The open questions are therefore practical as well as technical: whether the ranking holds on external centers, vendors, anatomies and acquisition protocols; how the metrics relate to clinically measured flow accuracy; and how sensitive the result is to different masks, hyperparameters and training budgets. The current paper does not answer those questions.
Access and disclosures
The authors report support from DFG Heisenberg, ERC StG EARTHWORM, ERC Proof-of-concept grant SYNCWORM and CAIMed, and declare no competing interests relevant to the article. The paper states that its source code is available through the listed GitHub repository.
Paper data and sources
Original title: FlowMoDL: Model-Based Deep Learning with Conjugate-Gradient Data Consistency for Highly Accelerated 4D Flow MRI Reconstruction
Authors: Tristan Gottwald, Michelle Bruch, Mubashir-Ul Hassan et al.
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-26
DOI: Not available
Original paper · Full text