{"id":"835092e2-957b-406a-b15d-05a05d3ab930","arxiv_id":"2502.08973","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A deep learning model estimates knee cartilage T1ρ maps from a proton-density-weighted FSE image and one T1ρ-weighted FSE image, with mean regional percentage error near 4% against an NLLS reference.","lead":"This paper tests whether a knee T1ρ relaxation map can be computed from one T1ρ-weighted MRI image plus a standard proton-density anatomical image, using deep learning. If it works, it could reduce scan time for cartilage health assessment in osteoarthritis patients.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Central claim depends on the U-Net learning a generalizable PD-to-TSL0 contrast correction, but Table 3's own NLLS baseline shows PD is not a valid I0; with only 40 subjects and single-scanner data, the sub-5% RPE is an in-distribution result, not a clinical-accuracy claim.","rationale":"The reader's weakest_assumption correctly identifies the load-bearing point: the feasibility claim rests on the network learning a stable correction for the PD-vs-TSL0 contrast mismatch that generalizes beyond the training distribution. My analysis agrees and sharpens the concern with a specific failure mode: the PD image is an uncalibrated surrogate for I0, and the NLLS baseline in Table 3 demonstrates that no simple signal-model relation holds. The deep network's success is therefore a learned mapping, and there is no evidence it is invariant to the acquisition normalization or scanner properties that determine c(x). The paper deserves credit for using five-fold cross-validation, reporting per-subject standard deviations, and explicitly listing limitations; the sub-5% RPE is a genuine in-distribution result. However, the clinical claim requires external validity, which is not demonstrated. The proposed intensity-scaling and gain-field test can be run on existing data and would settle whether the correction is physical or dataset-specific. I retain the reader's CONDITIONAL verdict: the method is promising but not yet clinically validated.","tokens_in":12667,"tokens_out":8399,"duration_ms":88896,"concrete_test":"Use the trained unmasked 2D U-Net from the PD/TSL=10 case and rerun inference on the held-out test folds after scaling the PD input by constants 0.8, 1.0, and 1.2, and after multiplying a test subject's PD volume by a smooth 2D gain gradient (e.g., +/- 15% across the image) to simulate coil shading. If RPE remains below 5% and MAE shifts by less than 1 ms for all perturbed inputs, the network is robust to arbitrary PD scaling and the central correction is physical; if either threshold is exceeded, the sub-5% RPE is tied to the training acquisition's normalization and will not generalize across sites.","verdict_should_be":"UNCHANGED","load_bearing_attack":"In Eq. (1), I0 is by definition the TSL=0 signal from the same magnetization-prepared acquisition. The PD-weighted FSE image P is not I0; at best P = c(x)*I0 with a scaling c that depends on coil sensitivity, fat suppression, registration, and resolution differences. The paper's own two-point NLLS reference (Table 3, PD-w + TSL=10ms) gives RPE 52.35 +/- 6.10%, quantifying how far P is from I0. The unmasked U-Net reduces this to 4.12 +/- 3.01% RPE, but with MAE 8.72 +/- 2.26ms and MAPE 19.00 +/- 3.39%: voxel-level errors remain large and cancel only in regional averaging. Since the network is trained and evaluated with 5-fold CV on 40 subjects from one 3T Philips system with a fixed PD protocol, the learned correction may be a dataset-specific mapping from P's intensity scale to the training T1rho distribution rather than the physical relation in Eq. (1). No external site, field strength, vendor, or changed PD normalization is tested. The conclusion that scan time is reduced is also not supported by Table 1: PD FSE is 7:20 min, while the T1rho row lists a 4:02 min acquisition; unless the PD volume is already acquired clinically, adding it to a T1rho exam increases total time.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript proposes an accelerated T1ρ quantification method for knee cartilage that uses one PD-weighted 3D FSE anatomical image and one T1ρ-weighted FSE image as inputs to deep learning models (a 2D U-Net and a 1D MLP), replacing the conventional acquisition of four T1ρ-weighted images. T1ρ maps estimated by these models are compared with ground truth maps from an NLLS fit of four T1ρ-weighted images (TSL=0,10,30,50 ms). The study evaluates four input combinations (PD-w or T1ρ-w TSL=0 with TSL=10 or 50 ms), three model variants, and masked versus unmasked U-Net inputs, using 5-fold cross-validation on 40 subjects from a single 3T Philips scanner. The authors report mean regional percentage error (RPE) below 5% for the best-performing models in all scenarios, including PD-based inputs, and concluded that the approach reduces scan time and maintains clinical standards.","tokens_in":13038,"tokens_out":5282,"duration_ms":48269,"significance":"If the claims hold, the approach could reduce the number of T1ρ contrast preparations needed for cartilage T1ρ mapping, potentially easing integration into clinical workflows. The experimental design is clear: 5-fold cross-validation, standard metrics (MAE, MAPE, RE, RPE), and ablations over inputs, architectures, and masking. The inclusion of an NLLS reference for each input combination is a useful benchmark. The central in-sample result—that a U-Net can map a PD image plus a single T1ρ-weighted image to a T1ρ map with mean regional error below 5%—is plausible and internally consistent. However, the clinical significance is currently overstated: the QIBA threshold is applied to a different statistical quantity, the scan-time reduction is not supported by Table 1, and no external validation is provided, so the sub-5% RPE remains an in-distribution result on a single-scanner, 40-subject dataset.","major_comments":[{"comment":"The translation of the QIBA within-subject coefficient of variation (4–5%) into a target RPE of 5% is not justified: the QIBA CV describes test–retest repeatability of repeated T1ρ measurements, not agreement between a learned estimate and an NLLS reference. Moreover, Table 3 reports mean RPE ± SD; for PD-w TSL=10 the mean RPE is 4.12 ± 3.01%, so a substantial fraction of subjects have RPE exceeding 5%. The abstract's 'below 5%' is a mean, not a per-subject guarantee, and the paper should report the percentage of subjects meeting the threshold.","section":"Section 2.4.5, Eq. (6)"},{"comment":"The hypothesis that a PD-weighted FSE image has contrast comparable to the TSL=0 T1ρ-weighted image contradicts the definition in Eq. (1), where I0 is the TSL=0 image from the same magnetization-prepared acquisition. The paper's own NLLS reference using PD as I0 gives RPE 52.35 ± 6.10% (Table 3), quantifying the contrast mismatch. The deep network therefore does not simply approximate Eq. (1) for PD-based inputs; it learns a dataset-specific mapping from PD to the T1ρ distribution. With 40 subjects from one 3T Philips system and a fixed PD protocol, the sub-5% RPE is an in-distribution result, and no external validation (different sites, field strengths, vendors, or PD normalization) is provided. The claim that the method 'maintains clinical standards' requires such validation.","section":"Section 2.1, Eq. (1), Table 3"},{"comment":"The claimed reduction in scan time is not supported by the acquisition parameters in Table 1. The PD-weighted FSE acquisition takes 7:20 min and the T1ρ-weighted acquisition takes 4:02 min; using both as inputs requires 11:22 min, which is longer than the four T1ρ-weighted images used for the ground truth (4:02 min). The conclusion that the method 'reduces scan time' is only valid if the PD volume is already acquired clinically as part of a standard knee MRI protocol, but the manuscript does not establish this condition. Please provide such evidence or revise the conclusion.","section":"Table 1, Abstract"},{"comment":"The reference NLLS method uses the same two inputs (e.g., PD and one T1ρ-weighted image) and incorrectly assumes that Eq. (1) holds with PD as I0. This is a physically invalid assumption, so the comparison in Table 3 does not demonstrate that the deep learning method 'outperforms' a legitimate NLLS approach; it demonstrates that a trained network can compensate for the contrast mismatch. A fairer NLLS baseline would fit with an unknown scaling factor between PD and I0, or the comparison should be framed as a learned mapping versus a physics-based fit under a misspecified model.","section":"Section 2.4.4, Table 3"},{"comment":"For PD-based inputs, the voxel-level errors remain large (e.g., unmasked U-Net with PD-w TSL=50: MAE 6.53 ± 1.04 ms, MAPE 14.73 ± 2.10%) while the RPE is 4.03 ± 2.63%. The low regional error results from cancellation of positive and negative voxel errors in regional averaging. The paper should report the full distribution of subject-level RPE (the boxplots in Figure 5 partially address this) and discuss whether the large voxel-level errors are clinically acceptable for applications that use voxel-wise T1ρ values, since the clinical claim is based on regional metrics.","section":"Section 3, Table 3"}],"minor_comments":[{"comment":"In the Discussion, 'a shorter T SKk' should read 'a shorter TSLk'.","section":"Section 4"},{"comment":"The caption contains a typo: 'osteoarthriths' should be 'osteoarthritis'.","section":"Figure 6 caption"},{"comment":"The limiter expression 'ˆy = {ymin, ReLU(x) + ymin, ymax}' is nonstandard; it should be written as a clamping operation, e.g., 'ˆy = clamp(ReLU(x) + ymin, ymin, ymax)', to avoid ambiguity.","section":"Equation (2)"},{"comment":"The column header 'PD-w2 TSL=10ms' is confusing because the superscript '2' appears to be a footnote marker but is not explained in the table caption. Please reformat the headers for clarity.","section":"Table 3"},{"comment":"The statement 'underwent Gaussian smoothing with a radius of three' does not specify the units of the radius (voxels or millimeters) or the kernel size; please clarify.","section":"Section 2.3.1"}],"recommendation":"major_revision","confidential_remarks":"The manuscript fits the scope of Magnetic Resonance in Medicine and the self-citations are appropriate for the dataset and sequence design. The in-sample experiments are carefully conducted, but the clinical and scan-time claims need to be substantiated or substantially tempered. I would like the editor to note that the QIBA threshold issue and the scan-time arithmetic are the two most important points for the revision; if they cannot be addressed with new experiments, the conclusions should be rewritten to reflect the current evidence."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nQuick take: this is a competent retrospective feasibility study, not a clinical breakthrough. The genuinely new bit is substituting a PD-weighted FSE image for the TSL=0 T1rho-weighted image and letting a U-Net correct the contrast mismatch. In this 40-subject, single-scanner dataset, the best models get regional percentage errors around 4% against the four-image NLLS reference, which is good evidence that the mapping can be learned.\n\nWhat I like: the experiments are well organized. Five-fold cross-validation with a fixed split, two architectures, an ablation on ROI masking, and per-subject standard deviations reported. The finding that masking hurts is non-obvious and worth reporting. The limitations section is honest, and the paper doesn't oversell the novelty — it explicitly builds on prior deep learning T1rho work.\n\nThe soft spots are real but not fatal. The PD image is not I0; the paper's own two-point NLLS with PD as I0 gives 52% RPE. So the U-Net is learning a contrast correction that is specific to this scanner, this PD protocol, and this training distribution. With 40 subjects and no external validation, sub-5% RPE is an in-distribution result. The QIBA repeatability coefficient cited for the 5% threshold is about test-retest reproducibility, not agreement with a numerical reference, so claiming 'clinical standards' is a category error. And the scan-time argument is shaky: the PD FSE takes 7:20 min while the T1rho series takes 4:02 min. Unless the PD is already part of the standard exam, you are adding time, not saving it. The savings only materialize in a protocol that already includes PD FSE.\n\nNone of this sinks the core feasibility claim. The method does appear to estimate regional T1rho values in this dataset. Who should read it: anyone working on accelerated quantitative MRI, especially T1rho/T2 mapping. It deserves a serious referee; I'd send it to review, keep the technical claims, and push the authors to temper the clinical language and, ideally, add a hold-out site or at least a different PD protocol.\n\nFor my own work: wouldn't cite it in the next year, but I'd bring it to reading group as an example of a clean feasibility study with a clear limitation in generalizability.","headline":"A clean feasibility study showing a U-Net can map PD-FSE plus one T1rho-weighted image to regional T1rho maps, but the clinical-accuracy claim overreaches and the scan-time savings depend on whether the PD is already in the protocol.","tokens_in":13565,"tokens_out":2501,"would_cite":false,"duration_ms":25473,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A deep network can estimate T1ρ maps of knee cartilage from a proton-density-weighted anatomical image and one T1ρ-weighted image, holding regional error below 5%.","keywords":["T1ρ mapping","knee cartilage","osteoarthritis","deep learning","U-Net","fast spin echo","proton density-weighted","quantitative MRI"],"falsifier":"Run the trained 2D U-Net on a held-out cohort scanned with a different scanner vendor or field strength; if the PD-weighted surrogate assumption is truly learned, the regional percentage error should stay below 5%, but if the network has overfit the 40-subject sampling, errors will rise sharply, mirroring the 52% RPE that the same PD-plus-TSL=10 input produces under NLLS fitting without learning.","tokens_in":12455,"feed_emoji":"🦵","tokens_out":9317,"duration_ms":86677,"temperature":0.7,"pith_summary":"This paper tries to establish that T1ρ relaxation maps of knee cartilage—a quantitative MRI biomarker for osteoarthritis—can be estimated from just two images: a proton-density-weighted anatomical fast spin-echo image plus a single T1ρ-weighted image, instead of the usual four T1ρ-weighted contrasts. The authors train a 2D U-Net and an MLP on 40 participants, using nonlinear least-squares fits from four T1ρ-weighted volumes as ground truth. The best deep-learning model in every tested input combination keeps regional percentage error below 5%, the clinical reproducibility target; for PD-plus-TSL=10 inputs the U-Net achieves 4.12 ± 3.01%, and for PD-plus-TSL=50 inputs 4.03 ± 2.63%. If this generalizes, it shortens knee cartilage T1ρ exams and lowers RF energy deposition, making quantitative cartilage assessment easier to fold into routine clinical MRI.","feed_headline":"Deep learning maps knee cartilage from two MRI images, under 5% error","feed_subtitle":"Using a PD-weighted anatomical scan plus one T1ρ-weighted image, a 2D U-Net matches conventional four-image fitting.","key_machinery":"The load-bearing object is a 2D U-Net with an output range limiter that estimates T1ρ voxel by voxel from a 64×64 patch pair of registered, Gaussian-smoothed PD-weighted and T1ρ-weighted FSE images. The limiter clamps predictions to [10, 100] ms using ReLU and a min-max gate, encoding prior knowledge that cartilage T1ρ lies in this range. Unlike an MLP that sees only voxel intensities, the U-Net pulls spatial context from both images; the paper's ROI-masking experiment shows that context outside the cartilage matters. The fitting target is the two-parameter mono-exponential equation $I_k = I_0 e^{-TSL_k/T_{1\\rho}}$, with the PD image standing in for $I_0$.","core_discovery":"The central claim is that a proton-density-weighted FSE image can replace the TSL=0 image in the standard two-parameter mono-exponential T1ρ model, provided a deep network performs the fit. The network is treated as a noise-robust universal approximator that learns to correct the contrast mismatch between PD-weighted and true TSL=0 contrast. Using only a PD-weighted FSE volume and one T1ρ-weighted volume acquired at TSL=10 or 50 ms, the best 2D U-Net produces regional percentage errors of roughly 4%, comparable to its performance with two true T1ρ-weighted images, while the two-point NLLS baseline with the same PD-based inputs fails at 52% and 33% RPE. The authors therefore claim that the method reduces the number of T1ρ contrast preparations to one, cuts scan time, and permits shorter spin-lock times that ease hardware and SAR constraints.","pith_inferences":["Whether the PD-for-TSL=0 substitution is safe across scanner manufacturers, field strengths, and fat-suppression variants is untested; a cross-site study would decide whether the learned correction is a physical mapping or a cohort-specific shortcut.","If the PD-to-TSL=0 mapping is learnable, the same two-image pattern may extend to T2 mapping or other joints, since many protocols already include proton-density-weighted anatomical images.","The paper points toward deriving T1ρ without dedicated spin-lock preparation, but its experiments stop short of that; testing this would require replacing the T1ρ-weighted input with a CPMG-derived image.","The U-Net's advantage over NLLS in low-SNR conditions suggests it regularizes noise; quantifying that bias-variance trade-off against the theoretical precision limit for unbiased estimators is a natural next step."],"forward_implications":["Knee cartilage T1ρ maps can be acquired with one T1ρ-weighted volume plus the PD-weighted anatomical volume already used in many knee protocols, rather than four T1ρ-weighted volumes.","Shorter spin-lock times, such as TSL=10 ms, remain within the under-5% regional error target, which relaxes RF amplifier and SAR constraints.","Deep-learning fitting from PD-plus-single-TSL inputs substantially outperforms two-point NLLS fitting from the same inputs, so the claimed gain is not simply the signal equation.","The 2D U-Net benefits from anatomical context outside the cartilage ROI; masking out those voxels degrades performance."],"supporting_citations":[{"why":"Sets the 4-5% test-retest reproducibility threshold that the paper adopts as its regional percentage error target.","marker":"[2]"},{"why":"Universal approximation theorem used to justify treating the neural network as a fitting approximator of the T1ρ signal equation.","marker":"[11]"},{"why":"Provides the 40-subject in vivo dataset with four TSL volumes and a PD-weighted FSE volume.","marker":"[17]"},{"why":"Symmetric diffeomorphic registration method that aligns the PD-weighted and T1ρ-weighted volumes before network input.","marker":"[18]"},{"why":"U-Net architecture that the paper modifies with a regressor and output range limiter.","marker":"[19]"},{"why":"MLP architecture adapted for the 1D comparison model used in the experiments.","marker":"[7]"},{"why":"B1/B0-compensated spin-lock preparation that supports the mono-exponential signal model in Equation 1.","marker":"[16]"}],"fun_headline_variants":["Deep learning cuts knee T1ρ scan to one contrast prep","PD-weighted image replaces one T1ρ in deep-learning mapping","U-Net maps knee cartilage from two images under 5% error","One T1ρ plus PD scan yields 4% error in knee cartilage map","Deep learning reduces required T1ρ contrasts to one for knee"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The result rests on the assumption that a PD-weighted FSE image can stand in for the TSL=0 T1ρ-weighted image and that the network's learned correction for that contrast mismatch keeps working on new patients; the paper's own NLLS baseline shows that without the learned correction the substitution fails (RPE ≈ 52%).","fun_headline_variants_meta":{"raw":{"variants":["Deep learning cuts knee T1ρ scan to one contrast prep","PD-weighted image replaces one T1ρ in deep-learning mapping","U-Net maps knee cartilage from two images under 5% error","One T1ρ plus PD scan yields 4% error in knee cartilage map","Deep learning reduces required T1ρ contrasts to one for knee"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000233,"raw_usage":{"total_tokens":1575,"prompt_tokens":1112,"completion_tokens":463,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":728,"completion_tokens_details":{"reasoning_tokens":371}},"tokens_in":728,"tokens_out":463,"duration_ms":4679,"temperature":1.0,"reasoning_tokens":371,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T23:03:17.600995+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the trained 2D U-Net on a held-out cohort scanned with a different scanner vendor or field strength; if the PD-weighted surrogate assumption is truly learned, the regional percentage error should stay below 5%, but if the network has overfit the 40-subject sampling, errors will rise sharply, mirroring the 52% RPE that the same PD-plus-TSL=10 input produces under NLLS fitting without learning.","supporting_citations":[{"cited_title":"U-Net: Convolutional Networks for Biomedical Image Segmentation","cited_arxiv_id":null,"evidence_quote":"U-Net architecture that the paper modifies with a regressor and output range limiter."},{"cited_title":"The QIBA Proﬁle for MRI-based Compositional Imaging of Knee Cartilage","cited_arxiv_id":null,"evidence_quote":"Sets the 4-5% test-retest reproducibility threshold that the paper adopts as its regional percentage error target."},{"cited_title":"Mul- tilayer Feedforward Networks Are Universal Approxima- tors","cited_arxiv_id":null,"evidence_quote":"Universal approximation theorem used to justify treating the neural network as a fitting approximator of the T1ρ signal equation."},{"cited_title":"A Systematic Post-Processing Approach for Quantitative $T_{1\\rho}$ Imaging of Knee Articular Cartilage","cited_arxiv_id":"2409.12600","evidence_quote":"Provides the 40-subject in vivo dataset with four TSL volumes and a PD-weighted FSE volume."},{"cited_title":"B., Epstein C","cited_arxiv_id":null,"evidence_quote":"Symmetric diffeomorphic registration method that aligns the PD-weighted and T1ρ-weighted volumes before network input."},{"cited_title":"Cram´ er–Rao Bound-Informed Training of Neural Net- works for Quantitative MRI","cited_arxiv_id":null,"evidence_quote":"MLP architecture adapted for the 1D comparison model used in the experiments."},{"cited_title":"T., Borthakur Arijitt, Elliott Mark A., et al","cited_arxiv_id":null,"evidence_quote":"B1/B0-compensated spin-lock preparation that supports the mono-exponential signal model in Equation 1."}],"review_version":1}