{"id":"cef24f0e-c140-428b-900d-598b789a1b09","arxiv_id":"2505.03149","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"A low-rank diffeomorphic velocity model, learned unsupervised from k-space data, recovers cardiac and respiratory resolved 5D MRI from a six-minute free-breathing scan.","lead":"This paper introduces DMoCo, an algorithm that reconstructs motion-resolved 3D heart images from a single six-minute MRI scan taken while the patient breathes normally and without ECG gating. It models every heart and breathing phase as a smoothly deformed version of one reference image and learns the image template and the motion model directly from raw MRI data.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Pure-deformation model Eq. (5) cannot separate contrast drift from motion; the Gd wash-out in the 6-min free-breathing acquisition is a realistic source of phase-indexed intensity change that the XCAT phantom does not include.","rationale":"The reader's weakest assumption is the same one I find most load-bearing: the model in Eq. (5) is a pure geometric-deformation model, so any intensity or contrast variation across phases must be explained by spurious deformation. The in-vivo protocol has a realistic source of such variation (gadolinium wash-out over a six-minute acquisition), and the XCAT phantom is constructed with static intensities, so the paper's only quantitative validation cannot detect this failure mode. I considered other concerns, such as the single-noise-realization phantom comparison and the lack of error bars; those affect statistical confidence in the claimed improvement, but they do not threaten the correctness of the in-vivo interpretation the way the intensity-drift identifiability problem does. The paper is transparent about being a feasibility study and about several limitations, but it does not mention this particular confound. My recommendation is unchanged relative to the reader's CONDITIONAL verdict: the method is promising and the math is coherent, but acceptance should remain conditional on a test that separates intensity drift from deformation. A contrast-drift phantom experiment is a direct, low-cost check that would settle whether the concern actually lands or whether the static-template model is robust enough in practice.","tokens_in":17639,"tokens_out":5491,"duration_ms":65205,"concrete_test":"Extend the XCAT simulation of Sec. V-A to include a realistic contrast washout: keep all motion ground truth unchanged, but multiply the blood-pool intensity (or a smooth spatial region) by a slowly varying factor, e.g., 1.0 decaying to 0.85 exponentially over the 5:58 acquisition. Run DMoCo on these data and compare the estimated diaphragm displacement and deformation fields against the known ground truth, and also compare against the no-drift XCAT run. If the estimated RD shifts by more than 1 mm, or if the recovered template develops apparent motion in static regions, the pure-deformation model is absorbing contrast drift as motion and the in-vivo claims require a phase-dependent intensity correction before acceptance. Repeat with at least five noise realizations so the result is not a single-run artifact.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Eq. (5) imposes rho_t(r) = eta(phi_tau(t)(r)): every phase volume is exactly a single static template warped by a diffeomorphism. The data-consistency term in Eq. (11) compares A_tau eta(phi_tau(r)) with b_tau, so the only degrees of freedom available for representing a phase-indexed signal change are the deformations phi_tau; there is no phase-dependent intensity or contrast factor. The in-vivo data were acquired roughly 15-20 minutes after gadolinium injection and over a fixed 5:58 minute acquisition (Sec. IV-D and V-C). Blood-pool and myocardial T1/contrast can drift during that window, and the paper explicitly notes the low CNR of the 3D acquisition. Such drift would be absorbed into spurious deformation, biasing the template, the motion fields, the respiratory displacement estimate, and the volumetric measures in Fig. 7. The XCAT phantom has static intensities, so it cannot detect this confounding. The paper's own limitation paragraph (Sec. VII) lists low CNR, lack of fat suppression, and runtime, but does not address the intensity-versus-deformation identifiability problem. This makes the central claim of improved in-vivo recovery vulnerable in exactly the regime the method is intended for, even though the low-rank diffeomorphic flow construction is internally coherent and the phantom results are supportive.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces DMoCo, an unsupervised motion-compensated reconstruction algorithm for free-breathing, ungated 3D cardiac MRI. Each phase volume is modeled as a deformation of a static template, where the deformation is obtained by integrating a velocity field along a path in a two-dimensional phase space (cardiac and respiratory). The velocity field is represented with a low-rank model combining spatial basis functions and an MLP that maps phase variables to weights; a path-independence penalty further constrains the diffeomorphisms. The template and motion parameters are learned directly from k-space data via a data-consistency loss. Experiments on an XCAT phantom and a 2D annulus phantom show that low-rank velocity representations are more compact than low-rank deformation or image representations, and comparisons against a reimplemented MoCo-SToRM baseline and a motion-resolved method suggest improved recovery. In-vivo results on five patients provide qualitative comparisons with 2D cine and preliminary left-ventricular functional measures.","tokens_in":17874,"tokens_out":5491,"duration_ms":55425,"significance":"If the claims hold, the paper contributes a principled and compact motion model that could improve motion-compensated cardiac MRI in low-CNR, heavily undersampled settings. The mathematical formulation is coherent, the unsupervised formulation avoids the need for supervised training data, and the path-independence penalty is a novel regularization. The phantom experiments provide quantitative support for the core modeling claim, and the use of an XCAT phantom with realistic acquisition settings is a strength. However, the validation is at the feasibility level: the phantom comparisons use a single noise realization, the in-vivo evaluation is largely qualitative, and the comparison baseline is a reimplemented variant of MoCo-SToRM rather than the original method. The paper therefore offers a promising algorithmic idea whose practical advantage is not yet fully established.","major_comments":[{"comment":"The model assumes each phase volume is exactly a deformation of a static template with no phase-dependent intensity or contrast variation. The in-vivo data were acquired 15–20 minutes after gadolinium injection and over a fixed 5:58 minute scan, so blood-pool and myocardial contrast can drift within the acquisition. Under the deformation-only model of Eq. (5), any such contrast drift would be absorbed into spurious deformations, biasing the template, motion fields, respiratory displacement estimates, and the functional measures in Fig. 7. The limitation paragraph (Sec. VII) mentions low CNR, fat suppression, and runtime but does not address this intensity-versus-deformation identifiability problem. The authors should either justify that contrast is effectively stationary over the acquisition window or demonstrate with a phantom experiment that includes time-varying intensity that the model does not introduce significant bias.","section":"Sec. III-A (Eq. 5) and Sec. V-C"},{"comment":"The quantitative XCAT comparison appears to be based on a single noise realization; no error bars or repeated trials are reported for the PSNR, SSIM, or diaphragm displacement values (e.g., 26.06 mm vs. 27.33 mm ground truth for DMoCo, 21.69 mm for MoCo-SToRM). Some differences, such as PSNR values of 29.87 dB vs. 29.69 dB for a given view, are small relative to typical run-to-run variability in stochastic optimization. To support the claim of improved recovery, the authors should report mean and standard deviation over multiple noise realizations and, if feasible, a significance test.","section":"Sec. VI-C and Fig. 4"},{"comment":"The MoCo-SToRM baseline is reimplemented with a low-rank deformation model (Eq. 14) rather than using the original CNN-based MoCo-SToRM from [23]. This changes the baseline architecture and optimization landscape; the observed improvement over this reimplementation may reflect differences in network capacity, initialization, or hyperparameter tuning rather than the proposed low-rank velocity representation. The authors should compare against the original MoCo-SToRM implementation, or explicitly justify that the reimplementation is a faithful and competitive proxy.","section":"Sec. IV-E and Eq. (14)"},{"comment":"The loss in Eq. (11) involves several hyperparameters (lambda1, lambda2, lambda3, rank R, K=20, path-perturbation variance), and the paper states these were chosen by trial and error on one dataset and kept fixed for other datasets. No sensitivity analysis is provided in the main text or appendices. Since the central claim is that the additional constraints (path independence and low-rank velocity) improve results, the paper would benefit from showing that the performance gain over MoCo-SToRM is robust to reasonable variations in these parameters, particularly lambda1 and R.","section":"Sec. IV-B and Eq. (11)"},{"comment":"The low-rank comparison in Fig. 2 uses velocity tensors estimated from the reference images via the proposed model (Eq. 13). Thus, the finding that velocities are more low-rank than deformations or images may be partly a consequence of the parameterization used to estimate them, rather than an intrinsic property of the underlying motion. The authors should either use analytically known velocities from the annulus phantom (where the motion is generated by known circle contractions) or clarify that the comparison is between fitted representations, which weakens the suggestion that the low-rank velocity model is inherently more compact.","section":"Sec. VI-A and Fig. 2"},{"comment":"The in-vivo validation against 2D cine uses only five patients and reports functional measures without statistical quantification (e.g., correlation coefficients, Bland-Altman limits, or confidence intervals). The plots show a best-fit line through the origin but no measures of agreement or variability. Given the abstract claims improved recovery over current motion-resolved and motion-compensated algorithms, the in-vivo comparison should either include a quantitative comparison against the SOTA baselines or be scoped more modestly as a feasibility demonstration with qualitative support.","section":"Sec. VI-D and Fig. 7"}],"minor_comments":[{"comment":"The sentence 'This intuition has led to the representation of diffeomorphisms, which are often expressed as the endpoint of a flow' is awkwardly phrased and should be rewritten for clarity.","section":"Sec. II-B"},{"comment":"The notation phi^1(r) and phi^0(r) is used without explicit definition; the superscripts denote the integration endpoint and should be explained.","section":"Eqs. (2)-(3)"},{"comment":"The auto-encoder loss in Eq. (10) appears to have a subscript inconsistency: the decoder is written as F_{theta1} while the encoder is E_{theta2}, which is swapped relative to the text. Please check and correct the notation, and ensure the loss is fully specified.","section":"Sec. IV-A, Eq. (10)"},{"comment":"The phrase 'piecewise-smooth profile with abrupt jumps' is contradictory; it should be 'piecewise-constant with abrupt jumps' or 'smooth with abrupt transitions'.","section":"Sec. VI-C"},{"comment":"The formatting of PSNR values, e.g., '29.60±(0.38)', is unusual and inconsistent; the parenthetical standard deviation notation should be clarified or made uniform.","section":"Fig. 3 caption"}],"recommendation":"major_revision","confidential_remarks":"The paper is a direct extension of the authors' prior MoCo-SToRM work, with the novel contribution being the low-rank velocity representation and path-independence penalty. The novelty is clearly stated, and the methodology is sound, but the validation is preliminary and several load-bearing concerns need to be addressed. The comparison against a reimplemented baseline and the lack of multiple noise realizations could be seen as overclaiming the improvement over SOTA methods. These issues are fixable within the scope of a revision, so I recommend major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Candid take: the low-rank diffeomorphic flow idea is genuinely new and the phantom experiment supports it, but the in-vivo claim carries an identifiability risk the paper doesn't address.\n\nThe core idea is worth your time. Instead of representing deformation fields directly, the paper models a low-rank family of velocity tensors and integrates them along paths in phase space. The observation that a low-rank velocity model (rank 3) beats higher-rank displacement or image models is the kind of transferable insight that makes a method worth knowing even before clinical validation. The XCAT comparison is decent: same sampling pattern, noise added, and the diaphragm displacement estimate (26.06 mm vs 27.33 mm ground truth) beats MoCo-SToRM (21.69 mm). The path-independence penalty is a reasonable way to constrain an otherwise ill-posed registration problem, and the in-vivo images do look cleaner than the motion-resolved baseline.\n\nNow the soft spots, in proportion. The biggest one is not any of the usual n=5 complaints. Eq (5) forces every phase volume to be exactly one static template deformed by a diffeomorphism. The in-vivo data were acquired 15-20 minutes after gadolinium injection over a fixed six-minute window. Blood-pool signal can drift during that window, and the model has no phase-dependent intensity term, so contrast drift will be absorbed into spurious deformation. That biases the template, the motion fields, and the reported volumes. The XCAT phantom has static intensities, so it cannot detect this confound. The limitation section mentions low CNR, fat suppression, and runtime, but not this identifiability problem. This is a real gap for the in-vivo claim, though it doesn't touch the phantom comparison.\n\nThe other issues are more minor. Phantom numbers are single noise realizations with no error bars. Hyperparameters were tuned on one dataset. No code or data are released. The in-vivo comparison is qualitative, and the baselines are the authors' own re-implementations. All of this is fixable and standard for a feasibility paper.\n\nBottom line: the method is coherent, the central phantom result is supportive, and the authors are honest about what is preliminary. This deserves a serious referee. Ask for error bars on the phantom, a code release, and at least a discussion—better, an experiment—separating intensity changes from deformation. If the identifiability issue is addressed, this becomes a solid methods paper.","headline":"Novel low-rank diffeomorphic flow model for 5D cardiac MRI, with convincing phantom evidence but an unaddressed contrast-drift confound that weakens the in-vivo claim.","tokens_in":18514,"tokens_out":2467,"would_cite":true,"duration_ms":24231,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A low-rank diffeomorphic flow model, learned directly from undersampled k-space data, reconstructs free-breathing, ungated 3D cardiac MRI into cardiac- and respiratory-resolved cine from a six-minute scan.","keywords":["cardiac MRI","motion compensation","diffeomorphic registration","low-rank model","free-breathing MRI","self-supervised reconstruction","radial k-space","5D imaging"],"falsifier":"Simulate the same free-breathing radial acquisition on a numerical phantom with known motion, then add a slow blood-pool intensity drift (for instance a ten percent linear gain over six minutes) and check whether the recovered deformations and template shift to compensate; if they do, the model has confused contrast change with motion.","tokens_in":17307,"feed_emoji":"🫀","tokens_out":10897,"duration_ms":102992,"temperature":0.7,"pith_summary":"This paper claims that six minutes of free-breathing, ungated 3D cardiac MRI can be turned into a cardiac- and respiratory-resolved 3D movie without breath holds, ECG gating, or supervised training data. The key move is to represent every phase volume as a fixed static template deformed by a diffeomorphism, with the deformations obtained by integrating a phase-parametrized velocity field that is low-rank across phases. The template and motion parameters are learned directly from undersampled k-space by minimizing data consistency plus penalties for flow path independence and template smoothness. On a simulated phantom the method recovers diaphragm displacement as 26.06 millimeters against a 27.33-millimeter ground truth, and it reports smoother time profiles and better respiratory estimates than the comparison methods; on five patients the functional measures track 2D cine. If correct, this would make cardiac functional MRI available to patients who cannot hold their breath and would deliver deformation maps usable for strain analysis.","feed_headline":"Six-minute free-breathing MRI becomes resolved 3D heart movies","feed_subtitle":"A low-rank flow learned from k-space removes breath holds and ECG gating from cardiac MRI.","key_machinery":"The load-bearing object is the low-rank velocity tensor $v_\\tau(s,r)=P(r)M_\\kappa(s,\\tau)$, in which the spatial basis $P$ is shared by all phases and the phase-dependent weights are produced by a small multilayer perceptron. The deformation at phase $\\tau$ is the endpoint of the ODE $d\\phi_\\delta/d\\delta=\\langle v_{\\gamma(\\delta)}(\\phi_\\delta),\\,\\gamma'(\\delta)\\rangle$ integrated along a straight-line path $\\gamma$ from the template phase to $\\tau$, and an alternate randomized path $\\beta$ supplies the path-independence penalty $\\mathbb{E}_{D_\\tau}\\|\\psi_\\tau-\\varphi_\\tau\\|^2$. Because every deformation is built from the same velocity basis through overlapping integrals, the model couples all phases together and is therefore more constrained than a direct low-rank deformation model.","core_discovery":"The central claim is that tissue velocities, rather than images or deformations, are the right low-rank object for jointly modeling cardiac and respiratory motion. The paper writes the velocity tensor at phase $\\tau$ as $v_\\tau(s,r)=\\sum_{n=0}^R p_n(r)m_n(s,\\tau)$, represents the image at phase $\\tau$ as $\\rho_t(r)=\\eta(\\varphi_{\\tau(t)}(r))$, and computes each deformation $\\varphi_\\tau$ by integrating the velocity field along a straight path through the two-dimensional phase space from the template to $\\tau$. A path-independence penalty compares that deformation with the one obtained along a randomly perturbed path, which further constrains the family. The static template $\\eta$ and the spatial basis $P$ and MLP weights $\\kappa$ are all fitted from undersampled k-space by stochastic gradient descent, with no fully sampled reference. The paper argues that overlapping path integrals and the path penalty make this deformation family smoother and more constrained than directly parameterized deformations, and that is why its reconstructed time profiles are less jumpy and its respiratory displacement estimates are closer to ground truth.","pith_inferences":["Going beyond the paper, a direct test of the contrast-drift risk would be to simulate the same six-minute radial acquisition with a slow blood-pool intensity drift and check whether the recovered motion fields absorb the drift; the paper's model has no mechanism to distinguish intensity change from deformation.","The same phase-space flow machinery should extend to other quasi-periodic motions in MRI, such as abdominal compression or fetal motion, whenever a low-dimensional phase coordinate can be measured from navigators or self-gating.","The path-independence penalty may erase genuine route dependence if cardiac contraction exhibits hysteresis; a phantom with known path-dependent motion could reveal whether this regularization is too strong.","The roughly three-hour reconstruction time on one GPU is the main practical obstacle the paper concedes; faster optimization or learned initialization would be needed before routine clinical use."],"forward_implications":["Free-breathing, ungated acquisitions of roughly six minutes can yield cardiac- and respiratory-resolved 3D cine, removing the breath hold and ECG gating requirements of current cardiac functional MRI.","Because the velocity tensor needs a much lower rank than images or deformation fields, motion estimates should stay stable under heavy undersampling and low contrast-to-noise ratio.","The path-independence penalty suppresses oscillatory, nonphysical deformations and abrupt jumps in the temporal profiles, which is what the paper says separates it from direct deformation models.","The reconstruction directly outputs smooth deformation fields for each cardiac and respiratory phase, which are available for myocardial strain analysis even though that analysis is future work."],"supporting_citations":[{"why":"The motion-compensated baseline the paper must outperform; it models deformations directly instead of integrating velocities.","marker":"[23]"},{"why":"The authors' prior 5D motion-compensated reconstruction whose deformation model DMoCo constrains further; also the source of acquisition settings.","marker":"[26]"},{"why":"The large-deformation diffeomorphic registration framework, giving the velocity-flow ODE (Eq. 2) that the method integrates.","marker":"[28]"},{"why":"The geodesic-flow formulation used to compute diffeomorphisms by ODE integration and to invert them by running the ODE backward.","marker":"[29]"},{"why":"The binned motion-resolved reconstruction used for initialization and as a comparison method.","marker":"[30]"},{"why":"The spiral-phyllotaxis 3D radial trajectory that sets the k-space sampling in all experiments.","marker":"[31]"},{"why":"The numerical phantom with known cardiac and respiratory ground truth used for quantitative error and diaphragm-displacement comparisons.","marker":"[32]"},{"why":"A recent 5D motion-compensated approach that relies on ferumoxytol contrast, cited to motivate working in the low-contrast, highly accelerated setting.","marker":"[25]"}],"fun_headline_variants":["Low-rank flow model frees cardiac MRI from breath holds","Velocity low-rank trick enables motion-corrected 3D cardiac MRI","Learning tissue velocities from k-space removes ECG gating","Low-rank diffeomorphic flow boosts free-breathing cine MRI","Unsupervised motion model sharpens free-breathing heart MRI"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The method assumes every measured image is exactly one fixed static template seen through a smooth invertible deformation, so any time-varying intensity change—such as contrast washout over the six-minute scan—has no separate representation and would be silently absorbed into false motion.","fun_headline_variants_meta":{"raw":{"variants":["Low-rank flow model frees cardiac MRI from breath holds","Velocity low-rank trick enables motion-corrected 3D cardiac MRI","Learning tissue velocities from k-space removes ECG gating","Low-rank diffeomorphic flow boosts free-breathing cine MRI","Unsupervised motion model sharpens free-breathing heart MRI"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000596,"raw_usage":{"total_tokens":2781,"prompt_tokens":928,"completion_tokens":1853,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":544,"completion_tokens_details":{"reasoning_tokens":1768}},"tokens_in":544,"tokens_out":1853,"duration_ms":13291,"temperature":1.0,"reasoning_tokens":1768,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T00:00:20.637906+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate the same free-breathing radial acquisition on a numerical phantom with known motion, then add a slow blood-pool intensity drift (for instance a ten percent linear gain over six minutes) and check whether the recovered deformations and template shift to compensate; if they do, the model has confused contrast change with motion.","supporting_citations":[{"cited_title":"Torres, Sean B","cited_arxiv_id":null,"evidence_quote":"The motion-compensated baseline the paper must outperform; it models deformations directly instead of integrating velocities."},{"cited_title":"Johnson, and Peder E","cited_arxiv_id":null,"evidence_quote":"The authors' prior 5D motion-compensated reconstruction whose deformation model DMoCo constrains further; also the source of acquisition settings."},{"cited_title":"Sodickson, and Ricardo Otazo","cited_arxiv_id":null,"evidence_quote":"The binned motion-resolved reconstruction used for initialization and as a comparison method."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The spiral-phyllotaxis 3D radial trajectory that sets the k-space sampling in all experiments."},{"cited_title":"Multi-dynamic deep image prior for cardiac MRI","cited_arxiv_id":"2412.04639","evidence_quote":"A recent 5D motion-compensated approach that relies on ferumoxytol contrast, cited to motivate working in the low-contrast, highly accelerated setting."}],"review_version":1}