Pith. sign in

REVIEW 4 major objections 6 minor 2 references

Bayesian Insights into Exchange and Restriction in Gray Matter Diffusion MRI

T0 review · 4 major / 6 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read The paper shows that per-voxel Bayesian uncertainty, not just the best-fit value, decides whether gray-matter diffusion MRI microstructure estimates are trustworthy, and that many commonly reported exchange-time and soma measurements are…

desk verdict First systematic SANDIX degeneracy analysis under human protocols; the qualitative message holds, but the uncertainty-filtering claims need a calibration check. read the letter →

arxiv 2508.19478 v3 pith:E6DDB3VR submitted 2025-08-26 physics.med-ph eess.IV

classification physics.med-pheess.IV
keywords DiffusionMRIGraymattermicrostructureWaterexchangeBayesianinferenceParameterdegeneracyUncertaintyquantificationNEXImodelSANDIX
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper asks whether the parameters returned by two diffusion MRI models of gray matter—NEXI, which adds water exchange between neurites and extracellular space, and SANDIX, which also adds a soma compartment—can be trusted voxel by voxel. Using a Bayesian deep-learning inference method that produces full posterior distributions, the authors show that extracellular diffusivity $D_e$ and neurite signal fraction $f$ are consistently recovered, whereas exchange time $t_{ex}$, intra-neurite diffusivity $D_i$, soma radius $r_s$, and soma fraction $f_s$ are frequently biased, imprecise, or degenerate under a practical 45-minute human acquisition protocol with realistic noise. Their practical claim is that uncertainty measures and degeneracy flags should be used to filter out unreliable voxels before interpreting results; in filtered in vivo cortical voxels, mean exchange time was about 5.51 ms for NEXI and 7.23 ms for SANDIX, lower than the 10–50 ms range usually reported. The reader should care because most prior studies fit these models without reporting per-voxel confidence, so biological conclusions drawn from unfiltered estimates may not be reproducible.

What carries the argument

The argument is carried by Bayesian posterior estimation with simulation-trained normalizing flows: a neural density estimator learns the conditional distribution of model parameters given the powder-averaged signal, and each voxel's fit is summarized by the maximum a posteriori value, an uncertainty score (interquartile range of the 50% most probable samples), and a flag for multimodal or degenerate posteriors. The models being fitted are NEXI, a two-compartment Kärger-form exchange model with four parameters ($t_{ex}$, $D_i$, $D_e$, $f$), and SANDIX, which adds an impermeable sphere compartment with radius $r_s$ and fraction $f_s$ through the Gaussian phase approximation, for six parameters total. This machinery matters because degeneracy and bias cannot be seen from a single best-fit value; the full posterior is what distinguishes trustworthy from untrustworthy estimates.

What would settle it

Compute empirical coverage on the paper's 1000-simulation test set: for each uncertainty threshold (10%, 30%, and 50%) and each parameter, count the fraction of cases where the ground-truth value falls inside the posterior's reported 50% credible interval. If low-uncertainty voxels contain the truth at substantially less than the nominal rate, the claim that uncertainty-based filtering selects trustworthy estimates is falsified.

Watch

Extended reading notes

Core claim

The central discovery claimed is that reliability in NEXI and SANDIX is parameter-dependent: $D_e$ and $f$ are well constrained across protocols, while $t_{ex}$, $D_i$, $r_s$, and $f_s$ are often not, with degeneracies hiding as single broad peaks under noise. In simulations, MAP bias and posterior uncertainty track each other—low-uncertainty estimates sit near the ground truth—so the paper treats posterior interquartile range as a usable quality score and multimodal posterior shape as a degeneracy flag. Applying these to in vivo Connectom data, the paper finds that only 40.85% (NEXI) and 26.73% (SANDIX) of cortical voxels pass a 10% uncertainty threshold for exchange time, and only 3.45% and 0.06% pass for soma radius and soma fraction. From the surviving voxels, the paper reports faster cortical water exchange ($5.51$ ms and $7.23$ ms) than commonly cited, and argues that non-linear least squares estimates that hit parameter boundaries are less interpretable than uncertainty-filtered Bayesian estimates.

Load-bearing premise

The load-bearing premise is that the Bayesian tool's uncertainty values are calibrated, meaning low-uncertainty voxels really are the accurate ones; the paper relies on this filtering step but never tests whether the stated uncertainty percentages match how often the true value falls inside them.

Editorial extensions

If this is right

  • Prior cortical exchange-time estimates in the 10–50 ms range include many high-uncertainty voxels; filtering to the most reliable voxels gives mean exchange times near 5.5 ms (NEXI) and 7.2 ms (SANDIX), implying faster neurite-to-extracellular water exchange in human cortex than usually reported.
  • Soma radius and soma fraction should not be interpreted from Connectom-level data without uncertainty filtering, because fewer than 4% of cortical voxels pass the 10% uncertainty threshold for either parameter.
  • Denser sampling of b-values and diffusion times reduces degeneracies and posterior uncertainty in simulations, making acquisition design the main practical lever for making exchange time and soma radius identifiable in humans.
  • Uncertainty-based selection improves scan–rescan consistency, so group-level comparisons in disease studies would be more reproducible if restricted to voxels with low posterior uncertainty.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper does not measure whether its posterior intervals are calibrated; a direct extension would be to compute empirical coverage on the test simulations for the 10%, 30%, and 50% uncertainty thresholds, and to recalibrate the thresholds if coverage is off.
  • If the low filtered exchange times of about 5–7 ms survive calibration and longer diffusion-time sampling, they would suggest that clinically feasible protocols need to sample much shorter exchange-sensitive timescales rather than simply adding more directions.
  • The same posterior-uncertainty and degeneracy-filtering logic generalizes naturally to other multi-compartment diffusion and relaxation–diffusion models, where parameter trade-offs are at least as severe.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. This paper investigates the reliability of parameter estimates from two gray-matter diffusion MRI models that incorporate water exchange, NEXI and SANDIX, using the μGUIDE Bayesian inference framework. The authors simulate data under two acquisition protocols (an extensive ex vivo protocol and an in vivo 3T Connectom protocol), with and without Rician noise, and compare parameter recovery for 1000 test cases. They then apply the trained estimators to in vivo human data from four volunteers, including scan-rescan sessions. The central findings are that some parameters (extra-cellular diffusivity De and neurite signal fraction f) are estimated robustly, while others (exchange time tex, intra-neurite diffusivity Di, soma radius rs, and soma fraction fs) show high uncertainty and bias, particularly under realistic noise and the reduced Connectom protocol. The paper further claims that filtering voxels by posterior uncertainty (thresholds of 50%, 30%, and 10%) selects trustworthy estimates, and that μGUIDE provides more robust and interpretable results than conventional NLLS fitting.

Significance. If the uncertainty-filtering claim is validated, the paper would provide a practically important protocol- and parameter-dependent identifiability map for NEXI and SANDIX, and would strengthen the case for reporting uncertainties in diffusion MRI microstructure fitting. The study is of clear interest to the dMRI community, and the simulation framework—1000 held-out test cases, two protocols, two models, and open code—is a useful resource. The scan-rescan reproducibility analysis is a welcome addition, and validating the estimator with the same forward model used for training is appropriate for an identifiability study and is not circular. However, the central practical conclusion depends critically on the calibration of the learned posteriors, which is not established in the manuscript.

major comments (4)
  1. [§2.3, §2.4, Tables 1 and 3] The paper's practical claim—that filtering voxels by posterior uncertainty yields trustworthy parameter estimates—requires the learned μGUIDE posteriors to be calibrated: the reported uncertainty (e.g., the 50% credible interval width) must contain the true value with the nominal frequency, and low-uncertainty MAPs must have correspondingly low error. The manuscript never tests this. Section 3.1 only states qualitatively that low-uncertainty estimates 'tend to coincide' with low-bias estimates (Figures 3 and 4), with no coverage statistic, simulation-based calibration (SBC) rank test, or quantitative error-versus-uncertainty analysis. Consequently, the 10%/30%/50% thresholds used to compute reliable-voxel percentages and mean filtered estimates (e.g., tex = 5.51 ms for NEXI) are not shown to select accurate estimates. I request an explicit calibration analysis—for example, the empirical coverage of the 50% credible interval as a function of the reported uncertainty, or an SBC diagnostic—and a justification for the chosen thresholds in terms of a target accuracy level.
  2. [Figure 7, §3.2, §4.2] The degeneracy detection rule is not specified. The text states that multi-modal posteriors indicate degeneracy and that degenerate posteriors are flagged with a red dot (Figures 3–6), but the algorithm or criterion used to decide that a posterior is multi-modal is never defined. This makes the reported degeneracy rates (Tables 1 and 3), and the claim in Section 3.1 that noise 'hides' degeneracies, non-reproducible. Please specify the detection procedure (e.g., clustering of posterior samples, number of modes from a Gaussian mixture fit, a dip test, or a threshold on a multimodality index) and, ideally, report the sensitivity of the degeneracy percentages to the chosen criterion.
  3. [§2.3] The μGUIDE versus NLLS comparison in Figure 7 is confounded by asymmetric filtering: μGUIDE distributions are thresholded by posterior uncertainty (50%, 30%, 10%) while NLLS results are shown unfiltered. The paper concludes in Section 4.2 that 'µGUIDE provides more robust and interpretable estimates than traditional NLLS fitting', but this design does not separate the effect of the inference method from the effect of voxel filtering. I recommend presenting both methods either unfiltered or with an analogous quality filter applied to NLLS (e.g., based on parameter bound violations or fit residuals), or explicitly reframing the comparison as 'μGUIDE with uncertainty filtering versus unfiltered NLLS'.
  4. The definition and normalization of the 'uncertainty' measure are ambiguous. The text defines it as 'the interquartile range of the 50% most probable samples' (Section 2.3), which is not a standard definition; the subsequent text and the numerical thresholds (10%, 30%, 50%) imply a width relative to the prior range (e.g., ∼15 ms for tex with a [1,150] ms range), but this is never stated explicitly. Please clarify whether the quoted uncertainty is the width of the 50% highest posterior density interval, and explicitly state how it is normalized so that the percentages in Tables 2 and 4 are interpretable.
minor comments (6)
  1. The phrase 'interquartile range of the 50% most probable samples' should be replaced by a precise description, e.g., 'width of the 50% highest posterior density interval', and the normalization relative to the prior range should be stated.
  2. The caption of Supplementary Figure 10 says 'Fitting results for the NEXI model using NLLS', but the text refers to both NEXI and SANDIX; this caption should be corrected to 'SANDIX'.
  3. The statement in Section 4.1 that degeneracies are '2.5 times more likely' under the Connectom protocol is not supported by a table or explicit percentages in the main text; please provide the actual numbers.
  4. The definition of the absolute neurite fraction f = (1 − fs) · fi uses fi without prior definition; please define fi explicitly in the SANDIX section.
  5. The justification for adding Rician noise at median SNR 50 to both protocols is reasonable, but reporting the SNR at representative b-values for each protocol would make the noise impact clearer.
  6. There are several citation format issues, such as 'uhlReducingNEXIAcquisition2025' in Section 4.1 and inconsistent spacing in some references; please check the bibliography and in-text citations for consistency.

Circularity Check

0 steps flagged · score 2.0 of 10

No significant circularity: reliability claims are tested against simulated ground truth and an independent NLLS baseline; self-citations to μGUIDE and the NEXI protocol are not load-bearing.

full rationale

The paper's central claims concern the accuracy, precision, and degeneracy of NEXI and SANDIX parameter estimates. These are evaluated on 1000 test simulations with known ground-truth parameters (Section 3.1, Figures 3-4) and on in vivo data compared with NLLS (Section 3.2, Figure 7), so the main results do not reduce to the model definitions or to a fitted parameter renamed as a prediction. The forward models are taken from the literature as objects of study, not derived from the data being predicted. The authors' own μGUIDE framework (Jallais & Palombo, 2024) and the NEXI 3T Connectom protocol (Uhl et al., 2024) are self-citations, but they are used as an inference tool and an acquisition scheme; the estimator is retrained for each model and protocol, and its outputs are checked against simulated ground truth, so these citations are not load-bearing evidence for the reliability conclusions. The uncertainty thresholds of 50%, 30%, and 10% are operational definitions of 'reliable' voxels, and the paper's simulation trend that low-uncertainty MAP estimates fall near the diagonal provides qualitative support for the link to accuracy. However, no posterior coverage or simulation-based calibration test is reported, so the calibration of the learned posteriors is not quantitatively established; this is a validation gap and a correctness risk, not a circular step. No equation in the paper is shown to be equivalent to its own input by construction.

Assumptions & free parameters 3 free parameters · 5 assumptions · 0 invented entities

No new physical entities, forces, or conserved quantities are introduced. μGUIDE is an existing software framework (Jallais & Palombo, 2024), not an invented entity for this paper. The free parameters listed are methodological choices that directly shape the quantitative reliability percentages and mean estimates reported as results.

free parameters (3)
  • Uncertainty filtering thresholds = 50%, 30%, 10%
    Used in Figures 7-8 and Tables 2/4 to define 'reliable' voxels; thresholds are chosen post hoc and directly determine the reported mean estimates and reliability percentages (e.g., tex mean 5.51 ms at <10% threshold).
  • Degeneracy detection criterion = not specified
    The paper flags degenerate posteriors as 'red dots' without defining the quantitative rule for multimodality; this uncontrolled choice affects degeneracy percentages in Tables 1 and 3.
  • MLP embedding size = 14 (NEXI), 22 (SANDIX)
    Chosen after 'preliminary tests' on range 5-30; affects the fidelity of the approximate posterior and therefore all uncertainty and degeneracy estimates.
assumptions (5)
  • domain assumption NEXI and SANDIX forward models faithfully represent gray matter diffusion signal
    The reliability analysis uses these models to generate ground-truth simulations; the paper does not validate the models against histology or independent measurements (Section 4.4 acknowledges this scope limitation).
  • domain assumption Kärger barrier-limited exchange model applies to neurites in gray matter
    Both models assume well-mixed compartments with exchange described by a single residence time; violations such as restricted exchange or branching would alter degeneracy structure (Section 2.1).
  • domain assumption Gaussian Phase Approximation for spheres is accurate
    SANDIX uses GPA to model diffusion inside somas; GPA errors at relevant radii and gradient timings would bias soma radius estimates (Section 2.1.2).
  • domain assumption Uniform priors over biologically plausible ranges
    Parameters sampled uniformly in tex [1,150] ms, f,fs in [0,1], Di,De in [0.1,3] um2/ms, rs in [1,30] um; posteriors are prior-dependent and no sensitivity analysis is given (Section 2.4).
  • ad hoc to paper NPE posterior approximation is accurate and calibrated
    μGUIDE's normalizing flow is assumed to give faithful posteriors; the paper does not report coverage or calibration checks, yet uses uncertainty thresholds as ground truth for reliability (Section 2.3).

how reviews work

0 comments
Cite this review

Pith. "Pith review of Bayesian Insights into Exchange and Restriction in Gray Matter Diffusion MRI." pith.science (2026). https://pith.science/paper/E6DDB3VR

@misc{pith2026250819478,
  author       = {Pith},
  title        = {Pith review of: Bayesian Insights into Exchange and Restriction in Gray Matter Diffusion MRI},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/E6DDB3VR}},
  note         = {Machine review of arXiv:2508.19478}
}
abstract

Biophysical models in diffusion MRI (dMRI) hold promise for characterizing gray matter tissue microstructure. Yet, the reliability of their parameter estimates remains largely under-studied, especially in models that incorporate water exchange. In this study, we investigate the accuracy, precision, and presence of degeneracy of two recently proposed gray matter models, NEXI and SANDIX, using established acquisition protocols, on both simulated and \textit{in vivo} data. We employ $\mu$GUIDE, a Bayesian inference framework based on deep learning, to quantify parameter uncertainty and detect degeneracies, enabling a more interpretable assessment of model fits. Our results show that while some microstructural parameters, such as extra-cellular diffusivity and neurite signal fraction, are robustly estimated, others, including exchange time and soma radius, are often associated with high uncertainty and estimation bias, particularly under realistic noise conditions and reduced acquisition protocols. Comparison with non-linear least squares fitting highlights the critical advantage of uncertainty-aware methods: the ability to flag and filter out unreliable estimates. Together, these findings emphasize the need to report uncertainty and account for model degeneracies when interpreting model-based estimates. Our study advocates for the integration of probabilistic fitting approaches into imaging pipelines to improve reproducibility and biological interpretability.

Figures

Figures reproduced from arXiv: 2508.19478 by the authors.

Figure 1
Figure 1. Graphical representation of the considered gray matter biophysical models, with the following [PITH_FULL_IMAGE:figures/full_fig_p005_1.png] view at source ↗
Figure 2
Figure 2. Simulated signals of the NEXI (A & C) and SANDIX (B & D) models, obtained from exemplar [PITH_FULL_IMAGE:figures/full_fig_p012_2.png] view at source ↗
Figure 3
Figure 3. Fitting results for the NEXI model using [PITH_FULL_IMAGE:figures/full_fig_p013_3.png] view at source ↗
Figures from the paper (7 more)
Figure 4
Figure 4. Figure 4: Fitting results for the SANDIX model using [PITH_FULL_IMAGE:figures/full_fig_p014_4.png]
Figure 5
Figure 5. Figure 5: Fitting results of the NEXI model on in vivo data acquired using the 3T Connectom protocol, estimated using µGUIDE. For each model parameter, the MAP estimate and associated uncertainty are shown across a brain slice. Voxels exhibiting degenerate posterior distribution…
Figure 6
Figure 6. Figure 6: Fitting results of the SANDIX model on in vivo data acquired using the 3T Connectom protocol, estimated using µGUIDE. For each model parameter, the MAP estimate and associated uncertainty are shown across a brain slice. Voxels exhibiting degenerate posterior distributi…
Figure 7
Figure 7. Figure 7: Comparison of parameter estimates across cortical ribbons from all participants using [PITH_FULL_IMAGE:figures/full_fig_p017_7.png]
Figure 8
Figure 8. Figure 8: Comparison of parameter estimates across cortical ribbon from one participant on two sessions [PITH_FULL_IMAGE:figures/full_fig_p019_8.png]
Figure 9
Figure 9. Figure 9: Fitting results for the NEXI model using NLLS on 1000 test simulations. Results are shown [PITH_FULL_IMAGE:figures/full_fig_p024_9.png]
Figure 10
Figure 10. Figure 10: Fitting results for the NEXI model using NLLS on 1000 test simulations. Results are shown [PITH_FULL_IMAGE:figures/full_fig_p025_10.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

2 extracted references · 1 linked inside Pith

  1. [448]

    C., Dyrby, T

    https://doi.org/10.1002/mrm.21646 Alexander, D. C., Dyrby, T. B., Nilsson, M., & Zhang, H. (2017). Imaging brain microstructure with diffusion MRI: Practicality and applications. NMR in Biomedicine , 32(4). https://doi.org/10. 1002/nbm.3841 Andersson, J. L., & Sotiropoulos, S. N. (2016). An integrated approach to correction for off-resonance effects and s...

  2. [1486]

    https://doi.org/10.1016/j.neuroimage.2006.10.037 Jones, D., Alexander, D., Bowtell, R., Cercignani, M., Dell’Acqua, F., McHugh, D., Miller, K., Palombo, M., Parker, G., Rudrapatna, U., & Tax, C. (2018). Microstructural imaging of the human brain with a ‘super-scanner’: 10 key advantages of ultra-strong gradients for diffusion MRI. NeuroImage, 182, 8–38. h...

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.