{"id":"a8c662d5-92de-4c3b-a4ab-c74b5494f7c2","arxiv_id":"2602.08918","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A public 3D+t MRI dataset of nine volunteers' thighs under controlled pressure-cuff deformations, with undersampled dynamic k-space data and fully sampled validation images for reconstruction benchmarking.","lead":"This paper releases a public MRI dataset for testing dynamic volumetric (3D+t) reconstruction algorithms, with undersampled scans of deforming thigh muscle plus fully sampled reference images. The dataset is meant to fill a gap in the field, where no public in vivo 3D+t ground-truth data existed before.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Ground-truth validity of the fully sampled validation images is under-substantiated: 1D-projection binning does not guarantee 3D deformation equivalence, and no quantitative validation is reported.","rationale":"The reader's weakest_assumption correctly identifies the repeatability/consistency of the pressure-cycle deformation as the load-bearing condition for the validity of the fully sampled validation images as ground truth. I agree that this is the central concern. The paper's technical validation is qualitative and uses a 1D projection surrogate that is insensitive to deformation components orthogonal to the readout axis, so it cannot rule out 3D mismatches. Additionally, the different temporal characteristics of the validation (static holds) and dynamic (continuous change) acquisitions could introduce viscoelastic effects that break the assumed equivalence. This is a genuine correctness risk for any downstream method that uses the dataset as a benchmark. However, the dataset itself and the paper's transparent description of its limitations (e.g., quasi-static ground truth, only one deformation fully validated) make the resource potentially valuable if the concern is addressed in a future version. Therefore, the reader's CONDITIONAL verdict is appropriate: the paper should be published only if accompanied by quantitative validation across subjects and bins, or if the limitations are clearly stated as only the hybrid scan's plateau state can be rigorously validated. I do not see a reason to move the verdict to ACCEPT or REJECT; the concern is specific and testable, not fatal.","tokens_in":8469,"tokens_out":6115,"duration_ms":62889,"concrete_test":"Compute quantitative similarity metrics (e.g., SSIM and normalized RMSE) between the binned dynamic reconstructions and the corresponding fully sampled validation images for all nine volunteers and all nine bins, not just the single representative volunteer shown. If the median SSIM falls below ~0.9, or if errors correlate with pressure level, the claimed consistency between dynamic and validation data is not supported. As a control, also compute the same metrics between the time-averaged dynamic image and the validation images to confirm that the binned images substantially outperform that baseline.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The dataset's utility as a validation resource hinges on the fully sampled validation images being valid ground truth for the undersampled dynamic data. This is asserted in Section 2.2.3 ('This pressure cycle resulted in repeatable elastic deformations') without quantitative evidence, and the Section 5 binning validation is only qualitative, using a 1D projection along the readout axis as the motion surrogate. A 1D projection collapses all information in the y-z plane, so two distinct 3D deformation states can yield identical projections; matching projections does not guarantee matching 3D anatomy. Moreover, the validation images were acquired after 22-second holds at constant pressure, while the dynamic scans used continuous pressure changes. Viscoelastic creep or pressure-rate effects can produce different deformation fields at the same nominal cuff pressure. Thus the 'high similarity' of the binned reconstructions to the validation images is not a rigorous test of the assumption that the dynamic data occupy the same deformation states as the validation data. Without this assumption, the dataset's fully sampled images may not be reliable as ground truth for dynamic reconstruction evaluation.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This data descriptor introduces a publicly available in vivo 3D+t MRI dataset for validating dynamic volumetric reconstruction algorithms. Nine healthy volunteers underwent controlled thigh deformations induced by a pneumatic pressure cuff. For each subject the authors provide multichannel undersampled k-space data from four dynamic scans, one of which (dynamic scan 1) is a six-repetition pressure cycle with an accompanying fully sampled validation scan at nine discrete pressure levels. The dataset also includes anatomical reference images, DIXON-based muscle segmentations, coil sensitivity maps, and noise measurements, all stored in ISMRMRD format under a Zenodo DOI. To illustrate use, the authors bin the six repetitions of dynamic scan 1 against the validation projections and report that the resulting reconstructions show 'high similarity' to the corresponding fully sampled validation images.","tokens_in":8704,"tokens_out":2933,"duration_ms":34850,"significance":"If the dataset is as described and the fully sampled images genuinely represent the deformation states of the dynamic data, this would fill a concrete gap: public 3D+t k-space data with ground-truth validation images are scarce. The provision of raw multichannel k-space, noise data for prewhitening, coil sensitivities, a code example, and a persistent DOI are clear strengths and lower the barrier for method development and fair benchmarking. The main weakness is that the validation of the central ground-truth assumption is qualitative and indirect; this is the load-bearing issue for any data descriptor that advertises its fully sampled images as validation references.","major_comments":[{"comment":"The ground-truth equivalence between the static validation images and the dynamic scan is asserted, not demonstrated. The validation images are acquired during 22-second holds at constant cuff pressure, whereas the dynamic scans use continuous inflation/deflation. Because soft tissue is viscoelastic, the deformation at a given nominal pressure may differ between steady-state holds and a continuously changing pressure cycle. The claim in Section 2.2.3 that 'This pressure cycle resulted in repeatable elastic deformations' needs quantitative support, e.g., per-repetition projection consistency, inter-repetition image similarity after binning, or a dedicated repeatability experiment. Without this, the fully sampled images may not be valid ground truth for the dynamic data.","section":"Section 2.2.2 vs. Section 2.2.3"},{"comment":"The technical validation is only qualitative ('high similarity') and relies on a 1D projection along the readout axis as the motion surrogate. This projection collapses all information in the y-z plane, so two distinct 3D deformation states can yield identical projections. The binning procedure therefore does not by itself certify that the dynamic data occupy the same deformation states as the validation images. Please report quantitative similarity metrics (e.g., SSIM, NRMSE, or mask overlap) between the binned reconstructions and the corresponding validation images, and ideally also show that the surrogate separates the nine bins in a meaningful way. This is important because the binned reconstructions are grouped using the validation projections, so some similarity is expected by construction; an independent comparison is needed.","section":"Section 5"},{"comment":"The sentence 'confirming that the induced deformations during the dynamic scans were consistent with those observed in the validation scan' goes beyond what the presented evidence supports. The qualitative similarity of binned images could also be explained by the binning process itself, since bins are formed by correlation to the validation projections. A more cautious conclusion, or a dedicated analysis of deformation-field correspondence (e.g., registration of the binned images to the validation images), would be appropriate.","section":"Section 5, last sentence"}],"minor_comments":[{"comment":"The text says 'The isometric knee flexion task of dynamic scan 3 was repeated with the rotated pressure cuff.' Based on the preceding items, this should likely be 'of dynamic scan 2,' since dynamic scan 3 already includes the rotated cuff without contraction. Please verify and correct.","section":"Section 2.2.3, item 4"},{"comment":"The statement that each bin is 'only slightly undersampled' could be made quantitative by stating the per-bin reduction factor or number of shots per bin. This would help readers assess the difficulty of the reconstruction problem.","section":"Section 5"},{"comment":"The figure caption mentions 'The first principal component of the projection data is used as motion surrogate signal,' but the main text does not explain how the principal component is derived from the projections. A brief sentence would improve reproducibility.","section":"Figure 5"},{"comment":"The advice to use rigid registration between validation and dynamic images is helpful, but the authors may also want to state whether the validation images are already in the same geometry as the dynamic scan 1 images or whether registration is always required.","section":"Section 6"}],"recommendation":"major_revision","confidential_remarks":"The dataset appears genuine and the technical infrastructure (ISMRMRD, raw k-space, noise data, code) is well aligned with community needs. The major concern is that the paper's central claim—that the fully sampled images serve as ground truth for the dynamic data—is supported only by a qualitative, 1D-projection-based comparison. This is fixable with additional analysis and more cautious wording, so I recommend major revision rather than rejection. If the authors can add quantitative metrics and a repeatability analysis, the paper could become a solid data descriptor."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Max,\n\nYou should know two things about this paper. First, it does fill a real gap: a public in vivo 3D+t MRI k-space dataset with fully sampled reference images. That is genuinely useful for people developing reconstruction algorithms. Second, the validity of those reference images as ground truth is under-substantiated. The authors bin undersampled dynamic data using 1D projections and then say the resulting images show 'high similarity' to the validation images, but they give no numbers, and the binning surrogate collapses 3D deformation information into a line.\n\nWhat the paper does well: the dataset is real (DOI in the abstract), the acquisition setup is simple and clearly described, the file format is standard (ISMRMRD), and they ship coil sensitivities, noise data, segmentations, and example code. The authors are explicit about what the data is and what it is not. That is exactly what a data descriptor should look like.\n\nThe soft spot is the ground truth itself. The fully sampled images were acquired at nine frozen pressure levels, each held for 22 seconds, while the dynamic scans used continuous inflation and deflation. Viscoelastic creep means tissue state at a given cuff pressure may differ between the two protocols. The paper claims the pressure cycle produced 'repeatable elastic deformations,' but no repeatability measurements are shown. And the validation uses a 1D readout projection as the motion surrogate, so two different 3D deformation states could produce identical projections. That means the binned images could look similar to the validation images without the underlying anatomy actually matching. This is a real gap, but it doesn't sink the dataset for all purposes—researchers can still use it as a shared testbed, provided they are careful about what they call ground truth.\n\nIf I were editing, I would send this to review, but I would want the authors to add quantitative similarity metrics (SSIM, NRMSE) for the binned reconstructions versus validation, and to discuss the static-versus-dynamic pressure issue. Ideally they would also show that the projection-based binning produces consistent results across repeated cycles.\n\nThe paper is for the MRI reconstruction community, specifically people working on 3D+t methods who need public in vivo data with reference images. It deserves a serious referee, and the dataset is likely to get used. I'd take it to our reading group if we were discussing data resources.","headline":"Useful new 3D+t MRI dataset with reference images, but the ground-truth validity is asserted on qualitative evidence; needs quantitative validation before being used as a benchmark.","tokens_in":9185,"tokens_out":2824,"would_cite":true,"duration_ms":29300,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":["87.61.-c","87.57.nj"],"model":"deepseek-v4-flash","headline":"This paper introduces a publicly available in vivo MRI dataset with fully sampled ground-truth images for validating 3D+t dynamic reconstruction methods.","keywords":["dynamic MRI","volumetric MRI","3D+t reconstruction","k-space dataset","validation data","muscle deformation","MRI reconstruction","in vivo"],"falsifier":"If someone acquires the dataset and performs a direct comparison between the binned dynamic reconstruction and the validation images, finding significant misalignment or intensity differences in the deformation region would disprove the repeatability assumption. Additionally, if the reported correlation between projections in the 1D motion surrogate is not consistently high across all subjects, that would indicate cycle-to-cycle variation.","tokens_in":8392,"feed_emoji":"🧲","tokens_out":950,"duration_ms":12710,"temperature":0.7,"pith_summary":"The paper's central claim is that a new, publicly released in vivo dataset fills a critical gap in dynamic volumetric MRI research: it provides undersampled multichannel k-space data alongside fully sampled validation images, enabling quantitative validation of 3D+t reconstruction algorithms. The authors establish this by describing a pneumatic-cuff setup that induces repeatable thigh-muscle deformations in nine healthy volunteers, with four dynamic scan types and a static validation scan at nine pressure levels. They further demonstrate a binning-based reconstruction that aligns six dynamic repetitions to the validation images, showing high similarity and thereby supporting the dataset's utility. The paper aims to give the research community a benchmark resource for developing and testing time-resolved volumetric MRI methods.","feed_headline":"New in vivo MRI dataset with ground truth for 3D+t reconstruction","feed_subtitle":"Nine volunteers, controlled muscle deformations, fully sampled validation data — a benchmark for dynamic volumetric MRI.","key_machinery":"The central mechanism is the pneumatic pressure cuff that produces repeatable, externally controlled muscle deformations, combined with a validation scan acquired at nine discrete pressure levels. The CASPR sampling pattern is used for dynamic acquisition, and a projection-based motion surrogate—derived from the center-line of k-space—enables binning the dynamic data into states matching the validation images. This binning approach is what validates the dataset's ability to provide ground-truth comparisons.","core_discovery":"The paper's contribution is the dataset itself: a publicly available collection of multichannel undersampled k-space data from nine healthy volunteers, acquired under controlled, repeatable deformations of the thigh muscles, with fully sampled validation images for one deformation type. The key discovery is that the induced pressure-cycle deformations are consistent enough across repetitions that data from six dynamic cycles can be binned into nine motion states and reconstructed to closely match the static validation images, confirming the feasibility of using this setup for validating 3D+t reconstruction methods.","pith_inferences":["The repeatability of the pressure-cuff-induced deformation is the load-bearing assumption; if deformations vary between cycles, the fully sampled validation images may not accurately represent the dynamic states, undermining the ground-truth validity.","The dataset could serve as a benchmark for motion-compensated reconstruction methods that aim to handle non-periodic or aperiodic motion, since the dynamic scans include variations like isometric knee flexion and cuff rotation.","By making both fully sampled and undersampled data available, the dataset allows for controlled studies of the trade-off between temporal resolution, spatial resolution, and reconstruction complexity.","The described binning approach itself—using k-space center projections for motion surrogate—could be extended to other dynamic imaging applications where periodic motion is assumed, providing a low-cost motion detection method."],"forward_implications":["The dataset enables quantitative validation of 3D+t MRI reconstruction methods, filling a gap where only 2D+t datasets with ground truth were publicly available.","Researchers can retrospectively undersample the fully sampled validation data to test reconstruction algorithms against true ground truth.","The controlled deformation setup supports studies of muscle strain, stiffness, and biomechanics, as well as dynamic image reconstruction.","The dataset's structure (multichannel raw k-space, coil sensitivities, noise measurements) allows development of advanced reconstruction methods like parallel imaging, compressed sensing, and low-rank techniques.","The provided segmentation masks and anatomical reference enable motion correction and registration-based validation workflows."],"fun_headline_variants":["MRI ground-truth dataset for dynamic 3D+t reconstruction","In vivo MRI data with ground truth for motion reconstruction","Benchmark dataset for dynamic volumetric MRI validation","Controlled muscle motion MRI dataset with ground truth","Fully sampled MRI data for validating 3D+t methods"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The pressure cycle is repeatable enough across the six repetitions that the fully sampled validation images at each pressure level are representative of the actual deformations in the dynamic data.","fun_headline_variants_meta":{"raw":{"variants":["MRI ground-truth dataset for dynamic 3D+t reconstruction","In vivo MRI data with ground truth for motion reconstruction","Benchmark dataset for dynamic volumetric MRI validation","Controlled muscle motion MRI dataset with ground truth","Fully sampled MRI data for validating 3D+t methods"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000506,"raw_usage":{"total_tokens":2286,"prompt_tokens":709,"completion_tokens":1577,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":453,"completion_tokens_details":{"reasoning_tokens":1500}},"tokens_in":453,"tokens_out":1577,"duration_ms":11960,"temperature":1.0,"reasoning_tokens":1500,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T03:03:58.766296+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"If someone acquires the dataset and performs a direct comparison between the binned dynamic reconstruction and the validation images, finding significant misalignment or intensity differences in the deformation region would disprove the repeatability assumption. Additionally, if the reported correlation between projections in the 1D motion surrogate is not consistently high across all subjects, that would indicate cycle-to-cycle variation.","supporting_citations":[],"review_version":1}