{"id":"f1820a31-c14b-4c56-987a-b1f6a00fcaa3","arxiv_id":"2508.19478","paper_version":3,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Using a deep-learning Bayesian fitting framework, the authors show that exchange time, soma radius, and soma fraction estimates from NEXI and SANDIX models are frequently unreliable under realistic noise and acquisition protocols, while extracellular diffusivity and neurite fraction are robust.","lead":"This study tests how reliably two brain imaging models (NEXI and SANDIX) can measure water exchange and cell body size in gray matter using Bayesian inference, finding that exchange time and soma size estimates are often too uncertain to trust under practical MRI protocols. It matters because these models are increasingly used to study brain tissue in health and disease, and the results argue for reporting uncertainty alongside every estimate.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No posterior calibration test; uncertainty-filter thresholds (10/30/50%) are unsupported, so the key claim that filtering yields trustworthy estimates is not established.","rationale":"The reader's weakest assumption — that μGUIDE posterior uncertainties are calibrated — is exactly the load-bearing point. The paper's quantitative claims (Tables 2/4, Figure 7, mean tex after filtering) all depend on the uncertainty thresholds selecting genuinely accurate estimates. Without a coverage test, the 10/30/50% thresholds are arbitrary; the paper even acknowledges thresholds are arbitrary (reader's rationale). The qualitative scatter in Figures 3-4 is suggestive but not sufficient: one needs a statistic relating uncertainty magnitude to actual error. I considered whether the unspecified degeneracy detection rule is more fundamental; however, the degeneracy percentages are descriptive and the main recommendation is filtering by uncertainty, so calibration dominates. The μGUIDE vs NLLS comparison is imperfect (no filtering in NLLS arm), but that is a secondary interpretation issue, not a threat to the internal validity of the uncertainty-filtering claim. The concrete SBC/coverage/error-decile checks would settle the matter and are fully within the authors' existing simulation framework. Thus the verdict should remain conditional: accept with the condition that calibration diagnostics be provided.","tokens_in":18637,"tokens_out":5718,"duration_ms":53713,"concrete_test":"Run simulation-based calibration (SBC) on the trained μGUIDE networks for each model/protocol/noise condition using the held-out test set: for each of the 1000 test signals, draw a sample from the learned posterior and compute the rank of the true parameter value in the posterior CDF; the distribution of ranks over all tests should be uniform. Additionally, compute empirical coverage of the reported 50% highest-posterior-density intervals: the fraction of test cases where the true parameter lies inside the interval should be ~50%. Finally, bin the test MAPs by uncertainty (e.g., deciles) and compare RMSE/bias between bins; the lowest-uncertainty bin should show substantially lower RMSE than the highest.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central practical claim is that voxels filtered by posterior uncertainty (Methods 2.3, Figure 7, Tables 2/4) yield reliable parameter estimates. This requires the learned μGUIDE posteriors to be calibrated: the 50% HPD width used as 'uncertainty' must actually contain the true value with the nominal frequency, and low-uncertainty MAPs must have low error. The paper never tests this. Section 3.1 only qualitatively states that 'low-uncertainty estimates tend to coincide with low-bias estimates' (Figures 3-4), without any coverage statistic, simulation-based calibration (SBC) rank test, or error-vs-uncertainty quantification. Scan-rescan reproducibility (Section 4.3) is not a substitute: a reproducible but miscalibrated posterior could consistently be overconfident. Consequently, the 10/30/50% thresholds used to compute 'reliable' voxel percentages and mean estimates (e.g., tex = 5.51 ms, rs clustered at ~10 µm) are not shown to select estimates with actual accuracy; if the posteriors are miscalibrated, the central conclusion that uncertainty-aware filtering yields trustworthy estimates is unsupported. The degeneracy detection rule is also unspecified, but the calibration gap is the more load-bearing issue.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper investigates the reliability of parameter estimates from two gray-matter diffusion MRI models that incorporate water exchange, NEXI and SANDIX, using the μGUIDE Bayesian inference framework. The authors simulate data under two acquisition protocols (an extensive ex vivo protocol and an in vivo 3T Connectom protocol), with and without Rician noise, and compare parameter recovery for 1000 test cases. They then apply the trained estimators to in vivo human data from four volunteers, including scan-rescan sessions. The central findings are that some parameters (extra-cellular diffusivity De and neurite signal fraction f) are estimated robustly, while others (exchange time tex, intra-neurite diffusivity Di, soma radius rs, and soma fraction fs) show high uncertainty and bias, particularly under realistic noise and the reduced Connectom protocol. The paper further claims that filtering voxels by posterior uncertainty (thresholds of 50%, 30%, and 10%) selects trustworthy estimates, and that μGUIDE provides more robust and interpretable results than conventional NLLS fitting.","tokens_in":18891,"tokens_out":5851,"duration_ms":49199,"significance":"If the uncertainty-filtering claim is validated, the paper would provide a practically important protocol- and parameter-dependent identifiability map for NEXI and SANDIX, and would strengthen the case for reporting uncertainties in diffusion MRI microstructure fitting. The study is of clear interest to the dMRI community, and the simulation framework—1000 held-out test cases, two protocols, two models, and open code—is a useful resource. The scan-rescan reproducibility analysis is a welcome addition, and validating the estimator with the same forward model used for training is appropriate for an identifiability study and is not circular. However, the central practical conclusion depends critically on the calibration of the learned posteriors, which is not established in the manuscript.","major_comments":[{"comment":"The paper's practical claim—that filtering voxels by posterior uncertainty yields trustworthy parameter estimates—requires the learned μGUIDE posteriors to be calibrated: the reported uncertainty (e.g., the 50% credible interval width) must contain the true value with the nominal frequency, and low-uncertainty MAPs must have correspondingly low error. The manuscript never tests this. Section 3.1 only states qualitatively that low-uncertainty estimates 'tend to coincide' with low-bias estimates (Figures 3 and 4), with no coverage statistic, simulation-based calibration (SBC) rank test, or quantitative error-versus-uncertainty analysis. Consequently, the 10%/30%/50% thresholds used to compute reliable-voxel percentages and mean filtered estimates (e.g., tex = 5.51 ms for NEXI) are not shown to select accurate estimates. I request an explicit calibration analysis—for example, the empirical coverage of the 50% credible interval as a function of the reported uncertainty, or an SBC diagnostic—and a justification for the chosen thresholds in terms of a target accuracy level.","section":"§2.3, §2.4, Tables 1 and 3"},{"comment":"The degeneracy detection rule is not specified. The text states that multi-modal posteriors indicate degeneracy and that degenerate posteriors are flagged with a red dot (Figures 3–6), but the algorithm or criterion used to decide that a posterior is multi-modal is never defined. This makes the reported degeneracy rates (Tables 1 and 3), and the claim in Section 3.1 that noise 'hides' degeneracies, non-reproducible. Please specify the detection procedure (e.g., clustering of posterior samples, number of modes from a Gaussian mixture fit, a dip test, or a threshold on a multimodality index) and, ideally, report the sensitivity of the degeneracy percentages to the chosen criterion.","section":"Figure 7, §3.2, §4.2"},{"comment":"The μGUIDE versus NLLS comparison in Figure 7 is confounded by asymmetric filtering: μGUIDE distributions are thresholded by posterior uncertainty (50%, 30%, 10%) while NLLS results are shown unfiltered. The paper concludes in Section 4.2 that 'µGUIDE provides more robust and interpretable estimates than traditional NLLS fitting', but this design does not separate the effect of the inference method from the effect of voxel filtering. I recommend presenting both methods either unfiltered or with an analogous quality filter applied to NLLS (e.g., based on parameter bound violations or fit residuals), or explicitly reframing the comparison as 'μGUIDE with uncertainty filtering versus unfiltered NLLS'.","section":"§2.3"},{"comment":"The definition and normalization of the 'uncertainty' measure are ambiguous. The text defines it as 'the interquartile range of the 50% most probable samples' (Section 2.3), which is not a standard definition; the subsequent text and the numerical thresholds (10%, 30%, 50%) imply a width relative to the prior range (e.g., ∼15 ms for tex with a [1,150] ms range), but this is never stated explicitly. Please clarify whether the quoted uncertainty is the width of the 50% highest posterior density interval, and explicitly state how it is normalized so that the percentages in Tables 2 and 4 are interpretable.","section":null}],"minor_comments":[{"comment":"The phrase 'interquartile range of the 50% most probable samples' should be replaced by a precise description, e.g., 'width of the 50% highest posterior density interval', and the normalization relative to the prior range should be stated.","section":null},{"comment":"The caption of Supplementary Figure 10 says 'Fitting results for the NEXI model using NLLS', but the text refers to both NEXI and SANDIX; this caption should be corrected to 'SANDIX'.","section":null},{"comment":"The statement in Section 4.1 that degeneracies are '2.5 times more likely' under the Connectom protocol is not supported by a table or explicit percentages in the main text; please provide the actual numbers.","section":null},{"comment":"The definition of the absolute neurite fraction f = (1 − fs) · fi uses fi without prior definition; please define fi explicitly in the SANDIX section.","section":null},{"comment":"The justification for adding Rician noise at median SNR 50 to both protocols is reasonable, but reporting the SNR at representative b-values for each protocol would make the noise impact clearer.","section":null},{"comment":"There are several citation format issues, such as 'uhlReducingNEXIAcquisition2025' in Section 4.1 and inconsistent spacing in some references; please check the bibliography and in-text citations for consistency.","section":null}],"recommendation":"major_revision","confidential_remarks":"The paper is a solid application of the authors' own μGUIDE framework to the NEXI/SANDIX models, and the central simulation results are likely to be of interest. The main blockers are methodological: the lack of a posterior calibration check, the unspecified degeneracy detection rule, and an unfair method comparison. These are fixable within the scope of a revision, so I recommend major revision rather than rejection. The self-citation to μGUIDE is substantial but appropriate given that the framework is the paper's own contribution."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This is a genuinely useful paper. It is the first systematic look at SANDIX identifiability under human acquisition protocols, using μGUIDE to quantify uncertainty and flag degeneracies, and it puts numbers on a suspicion the field has had: exchange time and soma radius are poorly constrained on a 3T Connectom protocol, while De and f come out reasonably stable. The simulations are well designed—1000 test cases, two protocols, noise-free and noisy—and the scan-rescan data add credibility. Code is public. The work deserves serious referee attention.\n\nWhat is genuinely new: previous NEXI studies already reported trouble estimating tex, but SANDIX's degeneracy structure under realistic protocols was largely unexplored. The head-to-head μGUIDE versus NLLS comparison on the same in vivo data, plus the reproducibility assessment, is a real addition.\n\nSoft spots, in order of size.\n\nFirst, the calibration gap. The paper's central practical claim is that filtering voxels by posterior uncertainty yields trustworthy estimates. But μGUIDE's posterior calibration is never tested. The simulations only qualitatively note that low-uncertainty MAPs tend to lie near the diagonal; there is no coverage statistic, no simulation-based calibration check, no error-versus-uncertainty curve. Scan-rescan reproducibility does not fill this gap: a consistently overconfident posterior is still reproducible. Since the 10/30/50% thresholds drive the reported percentages and the mean tex = 5.51 ms, those quantitative results rest on an unsupported assumption. This is fixable with a short SBC or coverage analysis, but without it the language about \"reliable voxels\" is too strong.\n\nSecond, the degeneracy detection rule is unspecified. The text says degenerate posteriors are flagged with a red dot, but never defines the criterion. Table 1's percentages depend on it. This is minor to moderate and easily fixed by reporting the rule.\n\nThird, the μGUIDE versus NLLS comparison filters μGUIDE's estimates by uncertainty while leaving NLLS unfiltered. That conflates the value of uncertainty information with the value of filtering. It doesn't sink the argument—NLLS intrinsically lacks such uncertainty measures—but the paper could acknowledge the asymmetry more sharply.\n\nBottom line: the qualitative message—do not trust tex, rs, fs under short human protocols—is likely correct and important. The quantitative reliability percentages need calibration support before I would cite them. I would accept conditional on adding a calibration check and specifying the degeneracy criterion. The paper is a solid, useful contribution that should shape how people interpret these models.","headline":"First systematic SANDIX degeneracy analysis under human protocols; the qualitative message holds, but the uncertainty-filtering claims need a calibration check.","tokens_in":19443,"tokens_out":3765,"would_cite":true,"duration_ms":35625,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper shows that per-voxel Bayesian uncertainty, not just the best-fit value, decides whether gray-matter diffusion MRI microstructure estimates are trustworthy, and that many commonly reported exchange-time and soma measurements are…","keywords":["Diffusion MRI","Gray matter microstructure","Water exchange","Bayesian inference","Parameter degeneracy","Uncertainty quantification","NEXI model","SANDIX model"],"falsifier":"Compute empirical coverage on the paper's 1000-simulation test set: for each uncertainty threshold (10%, 30%, and 50%) and each parameter, count the fraction of cases where the ground-truth value falls inside the posterior's reported 50% credible interval. If low-uncertainty voxels contain the truth at substantially less than the nominal rate, the claim that uncertainty-based filtering selects trustworthy estimates is falsified.","tokens_in":18457,"feed_emoji":"🧠","tokens_out":9524,"duration_ms":88932,"temperature":0.7,"pith_summary":"This paper asks whether the parameters returned by two diffusion MRI models of gray matter—NEXI, which adds water exchange between neurites and extracellular space, and SANDIX, which also adds a soma compartment—can be trusted voxel by voxel. Using a Bayesian deep-learning inference method that produces full posterior distributions, the authors show that extracellular diffusivity $D_e$ and neurite signal fraction $f$ are consistently recovered, whereas exchange time $t_{ex}$, intra-neurite diffusivity $D_i$, soma radius $r_s$, and soma fraction $f_s$ are frequently biased, imprecise, or degenerate under a practical 45-minute human acquisition protocol with realistic noise. Their practical claim is that uncertainty measures and degeneracy flags should be used to filter out unreliable voxels before interpreting results; in filtered in vivo cortical voxels, mean exchange time was about 5.51 ms for NEXI and 7.23 ms for SANDIX, lower than the 10–50 ms range usually reported. The reader should care because most prior studies fit these models without reporting per-voxel confidence, so biological conclusions drawn from unfiltered estimates may not be reproducible.","feed_headline":"Water-exchange estimates in cortex drop to ~5–7 ms when filtered","feed_subtitle":"Bayesian uncertainty filtering on diffusion MRI shows most exchange-time and soma-size voxels are unreliable.","key_machinery":"The argument is carried by Bayesian posterior estimation with simulation-trained normalizing flows: a neural density estimator learns the conditional distribution of model parameters given the powder-averaged signal, and each voxel's fit is summarized by the maximum a posteriori value, an uncertainty score (interquartile range of the 50% most probable samples), and a flag for multimodal or degenerate posteriors. The models being fitted are NEXI, a two-compartment Kärger-form exchange model with four parameters ($t_{ex}$, $D_i$, $D_e$, $f$), and SANDIX, which adds an impermeable sphere compartment with radius $r_s$ and fraction $f_s$ through the Gaussian phase approximation, for six parameters total. This machinery matters because degeneracy and bias cannot be seen from a single best-fit value; the full posterior is what distinguishes trustworthy from untrustworthy estimates.","core_discovery":"The central discovery claimed is that reliability in NEXI and SANDIX is parameter-dependent: $D_e$ and $f$ are well constrained across protocols, while $t_{ex}$, $D_i$, $r_s$, and $f_s$ are often not, with degeneracies hiding as single broad peaks under noise. In simulations, MAP bias and posterior uncertainty track each other—low-uncertainty estimates sit near the ground truth—so the paper treats posterior interquartile range as a usable quality score and multimodal posterior shape as a degeneracy flag. Applying these to in vivo Connectom data, the paper finds that only 40.85% (NEXI) and 26.73% (SANDIX) of cortical voxels pass a 10% uncertainty threshold for exchange time, and only 3.45% and 0.06% pass for soma radius and soma fraction. From the surviving voxels, the paper reports faster cortical water exchange ($5.51$ ms and $7.23$ ms) than commonly cited, and argues that non-linear least squares estimates that hit parameter boundaries are less interpretable than uncertainty-filtered Bayesian estimates.","pith_inferences":["The paper does not measure whether its posterior intervals are calibrated; a direct extension would be to compute empirical coverage on the test simulations for the 10%, 30%, and 50% uncertainty thresholds, and to recalibrate the thresholds if coverage is off.","If the low filtered exchange times of about 5–7 ms survive calibration and longer diffusion-time sampling, they would suggest that clinically feasible protocols need to sample much shorter exchange-sensitive timescales rather than simply adding more directions.","The same posterior-uncertainty and degeneracy-filtering logic generalizes naturally to other multi-compartment diffusion and relaxation–diffusion models, where parameter trade-offs are at least as severe."],"forward_implications":["Prior cortical exchange-time estimates in the 10–50 ms range include many high-uncertainty voxels; filtering to the most reliable voxels gives mean exchange times near 5.5 ms (NEXI) and 7.2 ms (SANDIX), implying faster neurite-to-extracellular water exchange in human cortex than usually reported.","Soma radius and soma fraction should not be interpreted from Connectom-level data without uncertainty filtering, because fewer than 4% of cortical voxels pass the 10% uncertainty threshold for either parameter.","Denser sampling of b-values and diffusion times reduces degeneracies and posterior uncertainty in simulations, making acquisition design the main practical lever for making exchange time and soma radius identifiable in humans.","Uncertainty-based selection improves scan–rescan consistency, so group-level comparisons in disease studies would be more reproducible if restricted to voxels with low posterior uncertainty."],"supporting_citations":[{"why":"Supplies the Bayesian posterior estimation framework (simulation-trained normalizing flow) that produces the uncertainty and degeneracy summaries central to the study.","marker":"Jallais & Palombo, 2024"},{"why":"Defines the NEXI model and the parameters $t_{ex}$, $D_i$, $D_e$, and $f$ whose reliability is being tested.","marker":"Jelescu et al., 2022"},{"why":"Introduces the SANDIX/SMEX exchange-plus-soma model and the extensive ex vivo acquisition protocol used in the simulations.","marker":"Olesen et al., 2022"},{"why":"Provides the 3T Connectom acquisition protocol and the non-linear least squares fitting baseline with Rician correction used for comparison.","marker":"Uhl et al., 2024"},{"why":"Contributes the SANDI soma compartment (impermeable spheres) and the assumption that soma exchange is negligible at the diffusion times considered.","marker":"Palombo et al., 2020"},{"why":"Reports previous in vivo high-gradient cortical exchange and soma estimates used for comparison and for the observation that large soma radii are hard to constrain.","marker":"Dong et al., 2025"},{"why":"Underpins the generalized Kärger ODE used by SANDIX to handle arbitrary gradient waveforms.","marker":"Ning et al., 2018"}],"fun_headline_variants":["Bayesian filtering: most exchange-time and soma-size estimates unreliable","Filtered cortex exchange time ~5-7 ms; most voxels fail uncertainty check","Uncertainty filtering reveals faster cortical water exchange in dMRI","Only minority of cortex voxels pass uncertainty test for exchange time","Bayesian quality scores show most gray matter model estimates unreliable"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the Bayesian tool's uncertainty values are calibrated, meaning low-uncertainty voxels really are the accurate ones; the paper relies on this filtering step but never tests whether the stated uncertainty percentages match how often the true value falls inside them.","fun_headline_variants_meta":{"raw":{"variants":["Bayesian filtering: most exchange-time and soma-size estimates unreliable","Filtered cortex exchange time ~5-7 ms; most voxels fail uncertainty check","Uncertainty filtering reveals faster cortical water exchange in dMRI","Only minority of cortex voxels pass uncertainty test for exchange time","Bayesian quality scores show most gray matter model estimates unreliable"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000233,"raw_usage":{"total_tokens":1523,"prompt_tokens":1003,"completion_tokens":520,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":619,"completion_tokens_details":{"reasoning_tokens":430}},"tokens_in":619,"tokens_out":520,"duration_ms":5466,"temperature":1.0,"reasoning_tokens":430,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T16:52:47.639940+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compute empirical coverage on the paper's 1000-simulation test set: for each uncertainty threshold (10%, 30%, and 50%) and each parameter, count the fraction of cases where the ground-truth value falls inside the posterior's reported 50% credible interval. If low-uncertainty voxels contain the truth at substantially less than the nominal rate, the claim that uncertainty-based filtering selects trustworthy estimates is falsified.","supporting_citations":[],"review_version":1}