REVIEW 4 major objections 6 minor 2 references
Bayesian Insights into Exchange and Restriction in Gray Matter Diffusion MRI
T0 review · 4 major / 6 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read The paper shows that per-voxel Bayesian uncertainty, not just the best-fit value, decides whether gray-matter diffusion MRI microstructure estimates are trustworthy, and that many commonly reported exchange-time and soma measurements are…
desk verdict First systematic SANDIX degeneracy analysis under human protocols; the qualitative message holds, but the uncertainty-filtering claims need a calibration check. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The argument is carried by Bayesian posterior estimation with simulation-trained normalizing flows: a neural density estimator learns the conditional distribution of model parameters given the powder-averaged signal, and each voxel's fit is summarized by the maximum a posteriori value, an uncertainty score (interquartile range of the 50% most probable samples), and a flag for multimodal or degenerate posteriors. The models being fitted are NEXI, a two-compartment Kärger-form exchange model with four parameters ($t_{ex}$, $D_i$, $D_e$, $f$), and SANDIX, which adds an impermeable sphere compartment with radius $r_s$ and fraction $f_s$ through the Gaussian phase approximation, for six parameters total. This machinery matters because degeneracy and bias cannot be seen from a single best-fit value; the full posterior is what distinguishes trustworthy from untrustworthy estimates.
What would settle it
Compute empirical coverage on the paper's 1000-simulation test set: for each uncertainty threshold (10%, 30%, and 50%) and each parameter, count the fraction of cases where the ground-truth value falls inside the posterior's reported 50% credible interval. If low-uncertainty voxels contain the truth at substantially less than the nominal rate, the claim that uncertainty-based filtering selects trustworthy estimates is falsified.
Extended reading notes
Core claim
The central discovery claimed is that reliability in NEXI and SANDIX is parameter-dependent: $D_e$ and $f$ are well constrained across protocols, while $t_{ex}$, $D_i$, $r_s$, and $f_s$ are often not, with degeneracies hiding as single broad peaks under noise. In simulations, MAP bias and posterior uncertainty track each other—low-uncertainty estimates sit near the ground truth—so the paper treats posterior interquartile range as a usable quality score and multimodal posterior shape as a degeneracy flag. Applying these to in vivo Connectom data, the paper finds that only 40.85% (NEXI) and 26.73% (SANDIX) of cortical voxels pass a 10% uncertainty threshold for exchange time, and only 3.45% and 0.06% pass for soma radius and soma fraction. From the surviving voxels, the paper reports faster cortical water exchange ($5.51$ ms and $7.23$ ms) than commonly cited, and argues that non-linear least squares estimates that hit parameter boundaries are less interpretable than uncertainty-filtered Bayesian estimates.
Load-bearing premise
The load-bearing premise is that the Bayesian tool's uncertainty values are calibrated, meaning low-uncertainty voxels really are the accurate ones; the paper relies on this filtering step but never tests whether the stated uncertainty percentages match how often the true value falls inside them.
Editorial extensions
If this is right
- Prior cortical exchange-time estimates in the 10–50 ms range include many high-uncertainty voxels; filtering to the most reliable voxels gives mean exchange times near 5.5 ms (NEXI) and 7.2 ms (SANDIX), implying faster neurite-to-extracellular water exchange in human cortex than usually reported.
- Soma radius and soma fraction should not be interpreted from Connectom-level data without uncertainty filtering, because fewer than 4% of cortical voxels pass the 10% uncertainty threshold for either parameter.
- Denser sampling of b-values and diffusion times reduces degeneracies and posterior uncertainty in simulations, making acquisition design the main practical lever for making exchange time and soma radius identifiable in humans.
- Uncertainty-based selection improves scan–rescan consistency, so group-level comparisons in disease studies would be more reproducible if restricted to voxels with low posterior uncertainty.
Reading between the lines
- The paper does not measure whether its posterior intervals are calibrated; a direct extension would be to compute empirical coverage on the test simulations for the 10%, 30%, and 50% uncertainty thresholds, and to recalibrate the thresholds if coverage is off.
- If the low filtered exchange times of about 5–7 ms survive calibration and longer diffusion-time sampling, they would suggest that clinically feasible protocols need to sample much shorter exchange-sensitive timescales rather than simply adding more directions.
- The same posterior-uncertainty and degeneracy-filtering logic generalizes naturally to other multi-compartment diffusion and relaxation–diffusion models, where parameter trade-offs are at least as severe.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper investigates the reliability of parameter estimates from two gray-matter diffusion MRI models that incorporate water exchange, NEXI and SANDIX, using the μGUIDE Bayesian inference framework. The authors simulate data under two acquisition protocols (an extensive ex vivo protocol and an in vivo 3T Connectom protocol), with and without Rician noise, and compare parameter recovery for 1000 test cases. They then apply the trained estimators to in vivo human data from four volunteers, including scan-rescan sessions. The central findings are that some parameters (extra-cellular diffusivity De and neurite signal fraction f) are estimated robustly, while others (exchange time tex, intra-neurite diffusivity Di, soma radius rs, and soma fraction fs) show high uncertainty and bias, particularly under realistic noise and the reduced Connectom protocol. The paper further claims that filtering voxels by posterior uncertainty (thresholds of 50%, 30%, and 10%) selects trustworthy estimates, and that μGUIDE provides more robust and interpretable results than conventional NLLS fitting.
Significance. If the uncertainty-filtering claim is validated, the paper would provide a practically important protocol- and parameter-dependent identifiability map for NEXI and SANDIX, and would strengthen the case for reporting uncertainties in diffusion MRI microstructure fitting. The study is of clear interest to the dMRI community, and the simulation framework—1000 held-out test cases, two protocols, two models, and open code—is a useful resource. The scan-rescan reproducibility analysis is a welcome addition, and validating the estimator with the same forward model used for training is appropriate for an identifiability study and is not circular. However, the central practical conclusion depends critically on the calibration of the learned posteriors, which is not established in the manuscript.
major comments (4)
- [§2.3, §2.4, Tables 1 and 3] The paper's practical claim—that filtering voxels by posterior uncertainty yields trustworthy parameter estimates—requires the learned μGUIDE posteriors to be calibrated: the reported uncertainty (e.g., the 50% credible interval width) must contain the true value with the nominal frequency, and low-uncertainty MAPs must have correspondingly low error. The manuscript never tests this. Section 3.1 only states qualitatively that low-uncertainty estimates 'tend to coincide' with low-bias estimates (Figures 3 and 4), with no coverage statistic, simulation-based calibration (SBC) rank test, or quantitative error-versus-uncertainty analysis. Consequently, the 10%/30%/50% thresholds used to compute reliable-voxel percentages and mean filtered estimates (e.g., tex = 5.51 ms for NEXI) are not shown to select accurate estimates. I request an explicit calibration analysis—for example, the empirical coverage of the 50% credible interval as a function of the reported uncertainty, or an SBC diagnostic—and a justification for the chosen thresholds in terms of a target accuracy level.
- [Figure 7, §3.2, §4.2] The degeneracy detection rule is not specified. The text states that multi-modal posteriors indicate degeneracy and that degenerate posteriors are flagged with a red dot (Figures 3–6), but the algorithm or criterion used to decide that a posterior is multi-modal is never defined. This makes the reported degeneracy rates (Tables 1 and 3), and the claim in Section 3.1 that noise 'hides' degeneracies, non-reproducible. Please specify the detection procedure (e.g., clustering of posterior samples, number of modes from a Gaussian mixture fit, a dip test, or a threshold on a multimodality index) and, ideally, report the sensitivity of the degeneracy percentages to the chosen criterion.
- [§2.3] The μGUIDE versus NLLS comparison in Figure 7 is confounded by asymmetric filtering: μGUIDE distributions are thresholded by posterior uncertainty (50%, 30%, 10%) while NLLS results are shown unfiltered. The paper concludes in Section 4.2 that 'µGUIDE provides more robust and interpretable estimates than traditional NLLS fitting', but this design does not separate the effect of the inference method from the effect of voxel filtering. I recommend presenting both methods either unfiltered or with an analogous quality filter applied to NLLS (e.g., based on parameter bound violations or fit residuals), or explicitly reframing the comparison as 'μGUIDE with uncertainty filtering versus unfiltered NLLS'.
- The definition and normalization of the 'uncertainty' measure are ambiguous. The text defines it as 'the interquartile range of the 50% most probable samples' (Section 2.3), which is not a standard definition; the subsequent text and the numerical thresholds (10%, 30%, 50%) imply a width relative to the prior range (e.g., ∼15 ms for tex with a [1,150] ms range), but this is never stated explicitly. Please clarify whether the quoted uncertainty is the width of the 50% highest posterior density interval, and explicitly state how it is normalized so that the percentages in Tables 2 and 4 are interpretable.
minor comments (6)
- The phrase 'interquartile range of the 50% most probable samples' should be replaced by a precise description, e.g., 'width of the 50% highest posterior density interval', and the normalization relative to the prior range should be stated.
- The caption of Supplementary Figure 10 says 'Fitting results for the NEXI model using NLLS', but the text refers to both NEXI and SANDIX; this caption should be corrected to 'SANDIX'.
- The statement in Section 4.1 that degeneracies are '2.5 times more likely' under the Connectom protocol is not supported by a table or explicit percentages in the main text; please provide the actual numbers.
- The definition of the absolute neurite fraction f = (1 − fs) · fi uses fi without prior definition; please define fi explicitly in the SANDIX section.
- The justification for adding Rician noise at median SNR 50 to both protocols is reasonable, but reporting the SNR at representative b-values for each protocol would make the noise impact clearer.
- There are several citation format issues, such as 'uhlReducingNEXIAcquisition2025' in Section 4.1 and inconsistent spacing in some references; please check the bibliography and in-text citations for consistency.
Circularity Check
No significant circularity: reliability claims are tested against simulated ground truth and an independent NLLS baseline; self-citations to μGUIDE and the NEXI protocol are not load-bearing.
full rationale
The paper's central claims concern the accuracy, precision, and degeneracy of NEXI and SANDIX parameter estimates. These are evaluated on 1000 test simulations with known ground-truth parameters (Section 3.1, Figures 3-4) and on in vivo data compared with NLLS (Section 3.2, Figure 7), so the main results do not reduce to the model definitions or to a fitted parameter renamed as a prediction. The forward models are taken from the literature as objects of study, not derived from the data being predicted. The authors' own μGUIDE framework (Jallais & Palombo, 2024) and the NEXI 3T Connectom protocol (Uhl et al., 2024) are self-citations, but they are used as an inference tool and an acquisition scheme; the estimator is retrained for each model and protocol, and its outputs are checked against simulated ground truth, so these citations are not load-bearing evidence for the reliability conclusions. The uncertainty thresholds of 50%, 30%, and 10% are operational definitions of 'reliable' voxels, and the paper's simulation trend that low-uncertainty MAP estimates fall near the diagonal provides qualitative support for the link to accuracy. However, no posterior coverage or simulation-based calibration test is reported, so the calibration of the learned posteriors is not quantitatively established; this is a validation gap and a correctness risk, not a circular step. No equation in the paper is shown to be equivalent to its own input by construction.
Assumptions & free parameters
free parameters (3)
- Uncertainty filtering thresholds =
50%, 30%, 10%
- Degeneracy detection criterion =
not specified
- MLP embedding size =
14 (NEXI), 22 (SANDIX)
assumptions (5)
- domain assumption NEXI and SANDIX forward models faithfully represent gray matter diffusion signal
- domain assumption Kärger barrier-limited exchange model applies to neurites in gray matter
- domain assumption Gaussian Phase Approximation for spheres is accurate
- domain assumption Uniform priors over biologically plausible ranges
- ad hoc to paper NPE posterior approximation is accurate and calibrated
Cite this review
Pith. "Pith review of Bayesian Insights into Exchange and Restriction in Gray Matter Diffusion MRI." pith.science (2026). https://pith.science/paper/E6DDB3VR
@misc{pith2026250819478,
author = {Pith},
title = {Pith review of: Bayesian Insights into Exchange and Restriction in Gray Matter Diffusion MRI},
year = {2026},
howpublished = {\url{https://pith.science/paper/E6DDB3VR}},
note = {Machine review of arXiv:2508.19478}
}
abstract
Biophysical models in diffusion MRI (dMRI) hold promise for characterizing gray matter tissue microstructure. Yet, the reliability of their parameter estimates remains largely under-studied, especially in models that incorporate water exchange. In this study, we investigate the accuracy, precision, and presence of degeneracy of two recently proposed gray matter models, NEXI and SANDIX, using established acquisition protocols, on both simulated and \textit{in vivo} data. We employ $\mu$GUIDE, a Bayesian inference framework based on deep learning, to quantify parameter uncertainty and detect degeneracies, enabling a more interpretable assessment of model fits. Our results show that while some microstructural parameters, such as extra-cellular diffusivity and neurite signal fraction, are robustly estimated, others, including exchange time and soma radius, are often associated with high uncertainty and estimation bias, particularly under realistic noise conditions and reduced acquisition protocols. Comparison with non-linear least squares fitting highlights the critical advantage of uncertainty-aware methods: the ability to flag and filter out unreliable estimates. Together, these findings emphasize the need to report uncertainty and account for model degeneracies when interpreting model-based estimates. Our study advocates for the integration of probabilistic fitting approaches into imaging pipelines to improve reproducibility and biological interpretability.
Figures
Figures from the paper (7 more)
Reference graph
Works this paper leans on
-
[448]
https://doi.org/10.1002/mrm.21646 Alexander, D. C., Dyrby, T. B., Nilsson, M., & Zhang, H. (2017). Imaging brain microstructure with diffusion MRI: Practicality and applications. NMR in Biomedicine , 32(4). https://doi.org/10. 1002/nbm.3841 Andersson, J. L., & Sotiropoulos, S. N. (2016). An integrated approach to correction for off-resonance effects and s...
arXiv 2017
-
[1486]
https://doi.org/10.1016/j.neuroimage.2006.10.037 Jones, D., Alexander, D., Bowtell, R., Cercignani, M., Dell’Acqua, F., McHugh, D., Miller, K., Palombo, M., Parker, G., Rudrapatna, U., & Tax, C. (2018). Microstructural imaging of the human brain with a ‘super-scanner’: 10 key advantages of ultra-strong gradients for diffusion MRI. NeuroImage, 182, 8–38. h...
arXiv 2018
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.