{"id":"e287ee65-a2e2-4501-8001-6b510b5c442a","arxiv_id":"2608.10769","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"A joint affine and diffeomorphic LDDMM registration framework improves Dice overlap on IXI brain MRIs over sequential FLIRT+LDDMM and foundation models CARL and uniGradICON.","lead":"This paper presents an image-registration method that estimates a global affine transformation and a local diffeomorphic deformation at the same time, rather than in the usual two sequential steps. On brain MRI it achieves slightly higher anatomical overlap than a standard sequential pipeline and two deep-learning models, without needing any training data.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 3D head-to-head 'joint affine + refinement vs FLIRT + LDDMM' confounds jointness with affine-objective mismatch; without a same-SSD separate-affine control, the 'better starting point' claim is underdetermined.","rationale":"The reader's weakest assumption concerns the separability of affine and diffeomorphic components, which is indeed an explicit admitted limitation. I agree that if the diffeomorphic component absorbs affine modes, the affine-only comparisons and the 'better starting point' interpretation are endangered. However, I find a more directly testable, experimentally grounded concern: the 3D evidence for the central claim compares a joint optimization affine against FLIRT, and these two affines differ not only in jointness but also in the affine objective, initialization, and optimizer. The paper itself demonstrates that FLIRT's affine-only Dice is higher, so the reversal after LDDMM could be due to the SSD-selected affine being better aligned with the SSD LDDMM stage. The missing control is the authors' own successive mode applied to the 3D cohort. This concern does not invalidate the reported Dice numbers; it means the central claim is currently underdetermined. Proposition 2.6 is indeed unproved, but I did not find a definite algebraic error in the decoupling, and the empirical claim can be tested independently of that proof. Given these addressable gaps, the existing CONDITIONAL verdict remains appropriate; no change is needed.","tokens_in":28484,"tokens_out":21988,"duration_ms":257451,"concrete_test":"Add to Table 2 an arm 'ours affine-only + LDDMM': estimate the affine with the same FA/DA optimizer and SSD data term but with the diffeomorphic covector frozen and gamma = 1 throughout (the 3D analog of the successive mode in Section 3.2), then apply the identical LBFGS LDDMM refinement used for 'ours (refinement)' and FLIRT+LDDMM. Report its median Dice, per-subject wins against FLIRT+LDDMM, and the Wilcoxon p-value. If this arm's median Dice and win count match 'ours (refinement)' within about 0.002, the advantage attributed to joint optimization is not established; if it clearly underperforms, the jointness claim is supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim, stated in Section 1, is that joint optimization produces an affine registration that is a better starting point for diffeomorphic matching than one estimated separately. The only direct 3D evidence is Table 2: 'ours (refinement)' beats FLIRT+LDDMM on 53/54 subjects while holding the final LDDMM stage fixed. But the two arms differ in two ways simultaneously: the affine comes from the joint Demeter/SSD optimization versus the separate FLIRT pipeline, and the affine-only objectives are not the same. Since the LDDMM refinement stage uses SSD, the joint affine was selected under the same metric as the refinement, whereas FLIRT's affine was not. Table 2 itself shows that FLIRT's affine-only Dice is significantly higher (reverse p = 2.3e-3), so the two affines are not comparable on their own metric; the final reversal could be caused by metric consistency rather than by jointness. The paper's own toy experiments include a successive mode using the same affine model with the diffeomorphic component deactivated (Section 3.2), but no equivalent 3D Dice comparison is reported. Without an arm that estimates an affine separately using the same SSD objective and optimizer and then applies the same LDDMM refinement, the experiment cannot attribute the improvement to joint estimation. If such a control matches 'ours (refinement)', the central claim collapses to an affine-objective consistency effect, not joint optimization.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces a joint affine-diffeomorphic registration framework within LDDMM. Two variational models are proposed: Full Affine (FA), where the deformation is generated by affine-plus-diffeomorphic controls, and Decomposed Affine (DA), where the affine part is restricted to rotations, translations, and anisotropic scalings. The authors derive Hamiltonian/geodesic equations for both models, introduce a change of variables intended to decouple the affine and diffeomorphic components, and present a numerical implementation with a progressive covector activation scheme and a variational weighting schedule. The method is evaluated on 2D toy examples and on 54 IXI brain MR volumes registered to the MNI template, comparing against FLIRT+LDDMM, CARL, and uniGradICON. The central claim is that joint optimization yields an affine component that is a better starting point for diffeomorphic refinement than an affine estimated separately.","tokens_in":28835,"tokens_out":5097,"duration_ms":57394,"significance":"If the central claim holds, the paper is a useful contribution to classical, unsupervised registration: it provides a geodesic formulation in which affine and diffeomorphic components are estimated concurrently, avoids the double interpolation of sequential pipelines, and outperforms the tested baselines on the IXI brain MRI task. The empirical section is unusually careful: paired one-sided Wilcoxon tests with Holm-Bonferroni correction, effect sizes, per-subject wins, and a public implementation are all provided. However, the central claim is currently underdetermined by the experimental design: the head-to-head comparison of \"ours (refinement)\" with FLIRT+LDDMM varies both the origin of the affine and the affine objective, so the reported advantage cannot be attributed specifically to joint estimation. The theoretical derivation also has a nontrivial proof gap in Proposition 2.6, and the paper explicitly concedes that the diffeomorphic component is not theoretically prevented from reproducing affine modes. These issues are fixable but load-bearing.","major_comments":[{"comment":"The central claim that joint optimization produces an affine registration that is a better starting point for diffeomorphic matching is not isolateable from the data. The comparison \"ours (refinement)\" vs. FLIRT+LDDMM holds the LDDMM refinement stage fixed, but the two affines differ in two ways simultaneously: one is produced by the joint Demeter/SSD optimization and the other by the separate FLIRT pipeline, whose affine-only objective is not SSD. The paper itself notes in the paragraph after Table 2 that FLIRT attains a significantly higher affine-only Dice than ours (reverse p = 2.3e-3). This means the observed final reversal could be caused by metric consistency between the affine stage and the subsequent SSD-based LDDMM refinement, rather than by joint estimation per se. The toy experiments in Section 3.2 include a successive mode with the same affine model and the diffeomorphic component deactivated, but no equivalent 3D Dice comparison with a same-objective, separately estimated affine is reported. I request an additional control arm: estimate an affine separately with the same SSD objective and optimizer used in the joint stage, then apply the same LDDMM refinement. If that control matches \"ours (refinement)\", the claimed advantage of joint estimation over separate estimation collapses to an affine-objective consistency effect.","section":"Section 3.3, Table 2"},{"comment":"The paper concedes in the conclusion that \"our proposed method does not theoretically ensure that the diffeomorphic component cannot reproduce global affine deformations.\" This concession directly affects the interpretation of the affine-only Dice comparisons and the claim that the joint affine is a better starting point. The variational weighting schedule in Section 3.1 is designed to encourage the affine component to absorb global motion first, but it does not provide a quantitative measure of mode separation. Without such a measure, the reported \"affine only\" Dice and the visual decomposition in Figure 5 rest on the assumption that the final diffeomorphism contains little affine content. I ask the authors to either (a) provide a quantitative test of affine content in the estimated diffeomorphic component, e.g., projection of the final velocity field onto the space of affine vector fields, or (b) temper the central claim to what is actually supported: that the joint pipeline produces a competitive final registration and a better SSD-consistent affine for refinement than FLIRT, without asserting that the affine was estimated independently of the diffeomorphic mode.","section":"Section 4"},{"comment":"Proposition 2.6 is the mathematical foundation for the FA implementation, but its proof is one sentence: \"The proof is similar to the one given in Definition 2.5.\" The stated geodesic equations are not a trivial modification of the earlier system: the new equations for p_A and p_b contain integral terms involving d\\tilde v_t, a Jacobian factor, and the pulled-back kernel K_{\\tilde V}. These terms require a careful derivation by differentiating the Hamiltonian with respect to \\tilde A and \\tilde b through the definition \\tilde v_t(x) = \\tilde A_t^{-1} v_t(\\tilde A_t x + \\tilde b_t). Since the numerical method in Section 3.1 integrates exactly these equations, an incorrect sign or missing term would propagate into all FA experiments. Please expand the proof, or at minimum provide a supplementary derivation, and verify the signs and the domain of integration in Eq. (26).","section":"Section 2.2.2, Proposition 2.6"}],"minor_comments":[{"comment":"There is a formatting error in the displayed Hamiltonian: the term \"|v\\|^2_V\" should read \"\\|v\\|^2_V\" for consistency with the surrounding notation.","section":"Eq. (16)"},{"comment":"The caption contains an incomplete sentence: \"top row diffeo grid as unlike Fig. 3 we cannot isolate the diffeomorphism component, bottom row ...\". This should be rephrased as a full sentence, e.g., \"Top row: diffeomorphism grid; unlike Fig. 3, the diffeomorphic component cannot be isolated from the affine, so the grid shows the full transformation.\"","section":"Figure 4 caption"},{"comment":"The variational schedule has several hand-chosen parameters (w, nu, kernel widths, per-component learning rates, multi-start count). The paper gives a useful qualitative account of how they were chosen, but no sensitivity analysis is reported. A short robustness experiment, e.g., varying w and nu over a plausible range on a small subset of the IXI subjects, would strengthen the claim that the schedule is not over-tuned to the reported results.","section":"Section 3.1, Eq. (45)"},{"comment":"The change of variables \\tilde I(x) = I(\\tilde A x + \\tilde b) is called a decoupling, but the subsequent data term is still evaluated through the composition with (\\tilde A_1,\\tilde b_1). The notation \\tilde D^\\mathrm{FA}_\\gamma(\\tilde A_1,\\tilde b_1,\\tilde I_1) is introduced without explicitly writing the final expression; please spell out the resulting data term to avoid ambiguity about which image is compared to the target.","section":"Section 2.2.2, Eq. (21)"}],"recommendation":"major_revision","confidential_remarks":"The main empirical comparison is well executed and the paper is honest about the limitation that the diffeomorphic component is not guaranteed to stay free of affine modes. The missing same-SSD separate-affine control is the key barrier; if the authors add it and the result still favors joint estimation, the central claim will be substantially stronger. The theoretical proof gap in Proposition 2.6 should be fixed before publication, as the numerical method relies on these equations. The paper's classification as math.DG is somewhat optimistic given that the core contributions are numerical and empirical, but the application is appropriate for a mathematical image analysis venue."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Read 2608.10769. Bottom line: the paper is a real contribution to classical LDDMM, with public code, careful paired statistics, and an honest limitations section. But the headline empirical claim — that joint optimization gives an affine that is a better starting point for diffeomorphic refinement than a separately estimated one — is not established by the 3D experiment as designed.\n\nThe comparison confounds jointness with objective mismatch. FLIRT's affine was fit with its own objective; the LDDMM refinement uses SSD. The joint affine was selected under that same SSD objective. Table 2 shows FLIRT's affine-only Dice is significantly higher, yet the composed and refined results favor the joint arm. That could be because jointness helps, or because the affine was selected under the same metric as the refinement. The 2D toy experiments include a successive mode with the same model and diffeo deactivated, but no Dice numbers are reported. Without a 3D control that fits a separate SSD affine with the same optimizer and then applies the same LDDMM refinement, the 'better starting point' claim stays underdetermined. The paper's own statement that the diffeomorphic component could absorb affine modes reinforces this concern.\n\nWhat is actually new and good: the image-space adaptation of the MP25 semidirect-product framework, the DA direct-product model, the transformed-kernel geodesic equations in Prop. 2.6, and the variational gamma schedule. The evaluation on 54 manually inspected IXI subjects with paired Wilcoxon tests and Holm correction is solid for what it compares. Beating CARL and uniGradICON without training is notable, even if the absolute Dice gains over FLIRT+LDDMM are small (0.841 vs 0.835 median, 0.847 after refinement).\n\nSoft spots, in order: Prop. 2.6's proof is one line despite new integral and Jacobian terms — needs a real derivation. The evaluation is single-anatomy, single-target, single-modality. The refinement step is an LBFGS wrap that the paper calls an optimizer artifact, which is fine but should be reported more prominently. None of these are fatal; the central theoretical framework and the engineering are sound.\n\nWho is this for: LDDMM practitioners, anyone building semidirect-product geodesic shooting, and readers tracking classical-versus-deep registration. I would bring it to reading group and cite it for the FA/DA construction. My recommendation for review: send it to a serious referee. The missing same-objective separate-affine control and the incomplete proof of Prop. 2.6 are addressable. If the authors add that control and derive the geodesic equations properly, the paper would be a solid accept.","headline":"A genuine, narrowly scoped LDDMM contribution whose headline claim about joint affine estimation is underdetermined by the 3D experiment, but which deserves a serious referee if the key derivations and controls are tightened.","tokens_in":29379,"tokens_out":2590,"would_cite":true,"duration_ms":30139,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A joint optimization beats affine-then-diffeomorphic registration on brain MRIs.","keywords":["image registration","LDDMM","affine transformation","diffeomorphic deformation","geodesic shooting","semidirect product","brain MRI","unsupervised registration"],"falsifier":"Take one of the 54 brain-MRI registration pairs, run the joint optimization, and least-squares fit the final diffeomorphic velocity field to the affine family $x\\mapsto Mx+b$ over the image domain; if the fitted affine projection has magnitude comparable to the estimated affine motion itself on subjects where the affine stage looks poor, the claimed separation between components is refuted and the joint affine cannot be credited with the improvement.","tokens_in":28275,"feed_emoji":"🧠","tokens_out":13310,"duration_ms":121043,"temperature":0.7,"pith_summary":"Standard image registration aligns one scan to a template in two fixed steps: fit a global affine transformation, then run a local diffeomorphic warp on the already-aligned image. This paper argues that the fixed order is itself the flaw, because the affine stage can absorb local deformation and hand the diffeomorphic stage a biased residual. It proposes estimating both parts simultaneously inside the large-deformation diffeomorphic metric mapping (LDDMM) framework, as a single geodesic problem in which an affine group and a diffeomorphism group act jointly. On 54 brain MRIs registered to a standard template, the joint estimate reaches a median Dice overlap of 0.841, or 0.847 after a short refinement, against 0.835 for the sequential affine-then-LDDMM baseline, and it outperforms two deep-learning foundation models using no training data. The paper's central claim is that an affine transformation chosen jointly is a better starting point for diffeomorphic refinement than an affine chosen alone, even when the isolated affine scores slightly higher on affine-only overlap.","feed_headline":"A joint optimization beats affine-then-diffeomorphic registration","feed_subtitle":"Optimizing global and local warps together yields median Dice 0.841 and 0.847, above 0.835 for the sequential baseline.","key_machinery":"The carrying object is the semidirect product group $\\mathrm{Aff}(\\mathbb{R}^d)\\ltimes \\mathrm{Diff}$, acting on an augmented image space that carries the current affine frame alongside the image. From this the paper derives optimality conditions for joint affine-diffeomorphic trajectories, uses a change of variables that decouples the affine motion from the diffeomorphic image evolution, and stabilises the joint optimization with two mechanisms: progressive covector activation, which gradually unlocks more complex affine degrees of freedom, and a variational weighting schedule driven by a rate estimator of the affine term, which forces the affine component to settle before handing over to the diffeomorphic part.","core_discovery":"At the core of the paper is a reversal in what makes a good affine initialization. The authors construct two unsupervised variational models: a Full Affine model in which the affine part is an unrestricted $GL(d)\\ltimes \\mathbb{R}^d$ motion, and a Decomposed Affine model in which the affine part is factorised into rotation, translation, and anisotropic scaling. For each model they derive Hamiltonian geodesic equations via the maximum principle of optimal control, and introduce a change of variables that removes the affine controls from the image evolution equation, so the two components can be read off separately. In the brain MRI experiments the affine-only part of the joint solution scores slightly below the affine-only part of the sequential baseline (median Dice 0.728 versus 0.731), yet once the diffeomorphic part is applied the joint solution wins on 49 of 54 subjects, and after refinement on 53 of 54. The paper takes this as direct evidence that affine-only overlap is a poor predictor of final registration quality: an affine optimum can leave a residual that is expensive for a diffeomorphism to absorb, while a joint affine is selected to leave a cheap residual. On 2D synthetic targets, the joint scheme also avoids the fold-over and grid-crossing pathologies the sequential pipeline exhibits.","pith_inferences":["A natural next test is to enforce the separation the paper leaves implicit, for example by penalising the projection of the diffeomorphic vector field onto affine modes; if the affine component then becomes identifiable at the population level, the model could support studies of global versus local anatomical variability.","Because the framework returns a full continuous geodesic rather than a single warp, the same formulation could feed trajectory-based analyses, such as parallel transport, atlas construction, or statistics on initial momenta, for which sequential pipelines provide no trajectory.","A testable prediction follows from the paper's own logic: the joint advantage over sequential registration should shrink when the initial misalignment is already small, since a near-perfect affine pre-alignment leaves little distinction between a cheap and an expensive residual for the diffeomorphic stage.","The variational handover schedule, which weights the affine term strongly until its rate of improvement stalls and then decays the weight, is a generic mechanism that could be transplanted to any joint estimator in which one component is far more expressive than the other."],"forward_implications":["Joint affine-diffeomorphic LDDMM is a workable unsupervised alternative to sequential affine-then-nonlinear pipelines, producing a single geodesic trajectory rather than a composed endpoint map.","Affine-only Dice is not a reliable predictor of final registration quality: an affine optimum can be worse for diffeomorphic refinement than an affine deliberately chosen jointly.","On the 54-subject brain MRI task, the joint result beats the sequential baseline on 49 of 54 subjects (53 of 54 after refinement) and beats both deep-learning foundation-model baselines on 53 of 54 (54 of 54 after refinement), with no training data.","Even in a joint formulation, the optimization trajectory must be actively controlled so the affine component resolves global motion before the more expressive diffeomorphic component takes over.","The joint scheme avoids the pathological grid-crossing deformations that appear in sequential affine-then-LDDMM registration on synthetic targets."],"supporting_citations":[{"why":"Supplies the semidirect-product construction coupling a finite-dimensional Lie group with a diffeomorphism group that the paper adapts to image registration.","marker":"[MP25]"},{"why":"Defines the LDDMM geodesic-flow formulation of diffeomorphic image matching that both proposed models build on.","marker":"[Beg+05]"},{"why":"Provides the optimal-control maximum-principle machinery used to derive the Hamiltonian geodesic equations.","marker":"[Arg+15]"},{"why":"Supplies the affine registration tool used as the sequential baseline and as the affine-only comparator.","marker":"[Jen+02]"},{"why":"Deep-learning baseline whose affine and diffeomorphic parts stay separable, used to test the joint-registration claim.","marker":"[Gre+25]"},{"why":"Deep-learning foundation-model baseline whose warp entangles global and local parts, used as a full-deformation comparator.","marker":"[Tia+24]"},{"why":"Supplies the PyTorch LDDMM implementation that the paper forks and extends.","marker":"[FGG21]"},{"why":"Documents the rough affine optimization landscape that motivates the paper's progressive activation and multi-start strategy.","marker":"[JS01]"}],"fun_headline_variants":["Joint affine-diffeomorphic optimization beats sequential pipeline","Don't fix affine first: joint model wins in brain MRI","Joint affine-diffeomorphic beats two-step registration","Affine-only overlap misleads: joint optimization wins"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the affine and diffeomorphic components actually separate during the joint optimization; the paper explicitly concedes it does not prevent the diffeomorphic part from reproducing global affine motion, so if the warp absorbs affine modes early, the affine-only scores and the 'better starting point' interpretation lose their basis.","fun_headline_variants_meta":{"raw":{"variants":["Joint affine-diffeomorphic optimization beats sequential pipeline","Don't fix affine first: joint model wins in brain MRI","Joint affine-diffeomorphic beats two-step registration","Affine-only overlap misleads: joint optimization wins"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001168,"raw_usage":{"total_tokens":4903,"prompt_tokens":1086,"completion_tokens":3817,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":702,"completion_tokens_details":{"reasoning_tokens":3753}},"tokens_in":702,"tokens_out":3817,"duration_ms":29081,"temperature":1.0,"reasoning_tokens":3753,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T17:43:51.734135+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take one of the 54 brain-MRI registration pairs, run the joint optimization, and least-squares fit the final diffeomorphic velocity field to the affine family $x\\mapsto Mx+b$ over the image domain; if the fitted affine projection has magnitude comparable to the estimated affine motion itself on subjects where the affine stage looks poor, the claimed separation between components is refuted and the joint affine cannot be credited with the improvement.","supporting_citations":[],"review_version":1}