{"id":"57e324ce-8c09-4ee6-8a2e-6cd27006b969","arxiv_id":"2412.15740","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"The paper presents a novel three-part taxonomy of regularization methods for medical image registration and identifies transfer gaps and evaluation weaknesses.","lead":"This review maps the landscape of regularization methods in medical image registration, sorting them into three categories: model based, problem specific, and learned. It gives researchers a structured reference for choosing and transferring regularization strategies, and it highlights gaps where conventional methods have not yet been adapted to deep learning.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'first comprehensive taxonomy' claim rests on an un-auditable search and on category boundaries that blur whether guidance losses and hyperparameter-selection networks count as regularization.","rationale":"The reader's conditional verdict identifies the same load-bearing weakness: the comprehensiveness and boundary-principle of the taxonomy are the least secure supports for the paper's central claim. I add two concrete manifestations. First, the paper itself admits that guiding loss terms are, strictly speaking, not regularization, yet it then classifies such terms as regularization in multiple sections and tables; this is an internal inconsistency in the object being taxonomized. Second, the learned-regularization category includes methods that only learn the weight of a fixed model-based term, which does not clearly satisfy the category's definition of learning deformation properties from data. Both issues are about the taxonomy's systematicity, not about the quality of the individual method descriptions. The paper has real strengths: the tables are extensive, the explicit/implicit distinction is useful, and the transfer-gap analysis is actionable. But those strengths do not by themselves establish the 'first comprehensive taxonomy' claim without an auditable search and a cleaner treatment of boundary cases. Since the reader already reached CONDITIONAL on essentially these grounds, my read does not move the verdict; it sharpens the specific checks that would settle whether the concern lands.","tokens_in":50754,"tokens_out":5526,"duration_ms":55611,"concrete_test":"Run a PRISMA-style literature search with the same cutoff (October 2024) across Scopus, PubMed, Google Scholar, IEEE Xplore, and Web of Science, using the paper's stated keywords plus their obvious synonyms. From the retrieved set, extract all methods proposed as regularization or guidance for pairwise deformable registration. Then have two independent annotators assign each method to Section 3's three categories using only the definitions in the text, and also flag any method that does not fit any category. If the search surfaces a non-negligible set of omitted regularization methods, or if inter-annotator agreement is below 0.8 Cohen's kappa, or if more than 10% of methods are assigned to two or more categories, the 'comprehensive and systematic' claim is not supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that this review fills a gap by providing the first comprehensive, structured taxonomy of regularization methods. Two conditions must hold for that claim: (1) the literature coverage is complete enough to be called comprehensive, and (2) the three categories are principled, exhaustive, and mutually exclusive. Neither condition has independent support in the manuscript. The entire literature search is described in one sentence (Sec. 1): three databases, keyword groups, no inclusion/exclusion criteria, no screening protocol, no list of screened or excluded papers. That makes the completeness claim non-auditable. The category boundaries are also not consistently defined. Section 2 states that guiding loss terms are 'strictly speaking, not a type of regularization,' yet Section 3.1.3 and Table 1 count image-space cycle-consistency losses as model-based regularization, and Section 3.2.1 counts Dice overlap losses as problem-specific regularization. Category III is similarly heterogeneous: the test-time methods in Sec. 3.3.3 (HyperMorph, CIN-based approaches) do not learn deformation properties; they learn to modulate the weight of a fixed diffusion regularizer, placing them at the boundary between model-based hyperparameter selection and learned regularization. No decision rule is given for such hybrids. If a significant fraction of methods straddle categories or fall outside them, the taxonomy is not the systematic map the abstract promises.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This review paper proposes a taxonomy for regularization methods in medical image registration, organizing the literature into three categories: model-based regularization with prior assumptions, problem-specific regularization with prior data knowledge, and learned regularization that derives deformation properties from training data. The paper surveys methods from both conventional and learning-based registration, summarizes them in tables, discusses the transfer of techniques between these frameworks, and concludes with open challenges and future directions. The central claim is that this is the first comprehensive, structured review of regularization in this field.","tokens_in":50992,"tokens_out":3480,"duration_ms":30562,"significance":"If the taxonomy is accepted, the paper could serve as a standard reference for regularization in medical image registration, which is a key component of almost every registration algorithm. The review's strengths include its broad coverage of the literature, the effort to categorize methods along multiple dimensions (explicit/implicit/guidance, conventional/learning-based), and the concrete identification of gaps in method transfer and evaluation. The paper is also honest in stating a caveat about guidance losses, even though this caveat creates an internal inconsistency in the taxonomy. The value of the work depends on whether the category boundaries are principled and consistently applied, and on whether the literature coverage is reproducible enough to support the 'comprehensive' claim.","major_comments":[{"comment":"The literature search is described in a single sentence (\"We searched for papers in Scopus, PubMed and GoogleScholar using combinations of the keywords ...\"), with no inclusion/exclusion criteria, no screening protocol, no counts of records screened or included, and no date range specification beyond \"up to October 2024.\" The abstract and introduction claim this is the first comprehensive taxonomy, but the completeness of the literature coverage cannot be verified or replicated from this description. Please add a systematic and auditable search protocol, including exact query strings, inclusion/exclusion criteria, a flow diagram of the screening process, and a list or at least summary counts of excluded papers.","section":"Section 1"},{"comment":"The paper states that guiding loss terms are \"strictly speaking, not a type of regularization\" (Section 2), but then classifies cycle-consistency losses as model-based regularization (Section 3.1.3 and Table 1) and Dice overlap losses as problem-specific regularization (Section 3.2.1 and Table 2). This is an internal inconsistency in the taxonomy's boundary. The authors should provide a decision rule: either broaden the definition of regularization to explicitly include guidance losses, or place guidance methods in a separate category, and apply this rule consistently across all tables and sections.","section":"Section 2 vs. Sections 3.1.3 and 3.2.1"},{"comment":"The test-time regularization methods (e.g., Hoopes et al. 2021, Mok and Chung 2021a) do not learn deformation properties; they learn to modulate the weight of a fixed diffusion regularizer, often through hypernetworks or conditional instance normalization. The category III definition in Section 3 and the abstract state that learned regularization \"derives deformation properties\" from training data, which these methods do not directly do. Clarify whether learning the mapping from a hyperparameter to network weights counts as learned regularization, and if so, refine the definition; otherwise, move these methods to a separate subcategory of hyperparameter adaptation.","section":"Section 3.3.3"}],"minor_comments":[{"comment":"The abstract states that learned regularization \"automatically derive[s] deformation properties from the data,\" but the test-time regularization methods in Section 3.3.3 primarily learn the effect of a predefined regularizer's weight. Align the abstract's wording with the actual scope of Section 3.3.3.","section":"Abstract"},{"comment":"The entry for \"Großbröhmer, 2024\" omits the co-author Heinrich, while the text in Section 3.2.1 refers to \"Großbröhmer and Heinrich (2024).\" Please make the citation consistent.","section":"Table 2"},{"comment":"The references Keall, Joshi, Vedam, Siebers, Kini, Mohan (2005a) and (2005b) appear to be the same paper with identical authors and title but are listed twice. Please verify and merge or distinguish them appropriately.","section":"References"},{"comment":"Figure 6 caption is somewhat vague about the two approaches being illustrated (PCA models and autoencoder networks); please make the caption self-contained, as the current text relies heavily on the body to explain the figure.","section":"Figure 6"}],"recommendation":"major_revision","confidential_remarks":"This is a well-written and broad survey that fills a real gap in the literature, and the issues raised are fixable within the manuscript's scope. The main concern is that the 'comprehensive' claim needs a reproducible search methodology, and the taxonomy boundaries need to be made explicit and consistent. I do not see a fundamental flaw that would require rejection, but the current version is not yet suitable for publication as is."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: the three-way taxonomy — model-based, problem-specific, learned — is a legitimate new way to organize regularization in medical image registration, and this will probably become a standard reference. The 'comprehensive' claim outruns the methodology they describe, but the underlying survey is sound and deserves serious peer review.\n\nWhat is genuinely new: prior surveys (Viergever, Haskins, Fu, Chen) touch regularization only in passing. This review makes it the object of study, organizes a couple hundred methods into three categories, and splits learned regularization into local smoothing, learned deformation spaces, and test-time adaptation. The strongest material is the transfer analysis: the observation that global L2 smoothing dominates (14 of 21 Learn2Reg 2022 methods) and that sliding motion, local rigidity, and physics-inspired regularization remain mostly absent from learning-based frameworks is correct and actionable.\n\nCredit where due: the tables are dense but usable, the notation is consistent, and the authors are honest about evaluation limits, including the unreliability of negative-Jacobian fractions and Dice as quality surrogates. The few self-citations (Reithmeir et al.) appear only as examples in the learned-regularization section; nothing circular there.\n\nSoft spots, in proportion. First, the literature search is described in one sentence: three databases, keyword combinations, no inclusion/exclusion criteria, no screening flow, no list of excluded papers. For a headline claim of 'first comprehensive overview,' that is genuinely un-auditable. I do not think coverage is bad, but a referee should request a search appendix or a toned-down claim. Second, the guidance-loss boundary is fuzzy by the authors' own admission: Section 2 says guiding losses are 'strictly speaking, not a type of regularization,' yet Tables 1 and 2 count cycle-consistency and Dice losses as regularization. Transparent, but inconsistent with the abstract's promise of a systematic map. Third, the stress-test's objection to test-time methods lands only mildly. HyperMorph and the CIN methods modulate the weight of a fixed diffusion term rather than learning deformation properties, which puts them near hyperparameter selection, but the authors frame the category as 'what is learned,' and these methods do learn something useful. That is a boundary quibble, not a category error.\n\nVerdict: the central taxonomy holds up. The search protocol and the guidance-loss classification need work; the survey itself does not. A serious editor should send this to review, and a good referee should focus on those two points. After revision, I expect this to become the default citation for regularization in registration.","headline":"A genuinely useful taxonomy of regularization in medical image registration that will likely become a standard reference, but its 'comprehensive' claim needs a real search protocol or softer language.","tokens_in":51529,"tokens_out":4983,"would_cite":true,"duration_ms":41168,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"All regularization methods in medical image registration sort into three families — model based, problem specific, and learned — ordered by prior information, and the resulting map exposes the field's default to plain smoothing.","keywords":["medical image registration","regularization","learning-based registration","taxonomy","learned regularization","sliding motion","test-time regularization","deformation spaces"],"falsifier":"Run the review's own search protocol independently — the same three databases, the same keyword combinations, the same October 2024 cutoff — and compare the recovered method set against the review's tables; if a substantial body of published regularization methods is missing or cannot be assigned to exactly one of the three families, the claim of comprehensive structured coverage fails. A cheap first probe is to count how many methods cited within the review's own reference list never appear in its category tables.","tokens_in":50560,"feed_emoji":"🗺️","tokens_out":15423,"duration_ms":119365,"temperature":0.7,"pith_summary":"Medical image registration aligns two images by estimating the deformation between them, and because the problem is ill-posed — many deformations fit, but few are anatomically plausible — the regularization term is what makes the answer trustworthy. This review claims that prior surveys treated regularization as a side topic and supplies the first structured taxonomy of the field: model-based methods that impose global assumptions such as smoothness, invertibility, and diffeomorphism; problem-specific methods that inject segmentation maps, boundary locations, or physiological context; and learned methods that derive deformation properties directly from training data. The three families are ordered by the amount of prior information they encode, with problem-specific and learned methods growing out of model-based roots. If the taxonomy is right, it gives practitioners a map for transferring solutions across applications, gives the field a shared vocabulary, and documents why state-of-the-art frameworks still default to plain diffusion smoothing despite the richer alternatives available.","feed_headline":"Three families organize all image-registration regularization methods","feed_subtitle":"A structured map of conventional and learned techniques shows many modern methods still default to plain smoothing.","key_machinery":"The carrying mechanism is the taxonomy itself: three families — model based, problem specific, and learned — arranged on an ascending scale of how much prior information the regularization encodes. Two supporting distinctions do the classification work: explicit versus implicit regularization (a penalty term added to the objective versus smoothness that arises from the transformation model's parameterization, such as B-spline free-form deformations or multiresolution schemes), and guiding loss terms that indirectly enforce plausibility without operating on the deformation field, such as segmentation overlap measures. The Jacobian determinant $\\det\\mathbf{J}$ of the displacement field is the common technical language throughout: it quantifies folding ($\\det\\mathbf{J}<0$), volume change ($\\det\\mathbf{J}>1$ or $<1$), and volume preservation ($\\det\\mathbf{J}=1$), and most invertibility, incompressibility, and rigidity constraints are stated directly in terms of it. The subcategories inside each family — for example learned regularization split into learned local smoothing, learned deformation spaces, and learned test-time regularization — are what make the transfer analysis possible, since they allow the review to track which conventional methods have or have not been adapted into learning-based frameworks.","core_discovery":"The paper's central claim is that every regularization method proposed for pairwise medical image registration, conventional or learning-based, can be placed into exactly one of three families organized by the source and amount of prior information, with the literature covered through October 2024. Model-based regularization applies a user-defined assumption globally — smoothness through diffusion or curvature penalties, invertibility through constraints on the Jacobian determinant $\\det\\mathbf{J}$, inverse- or cycle-consistency, diffeomorphic parameterization, volume preservation, and physics-inspired elasticity or viscous-fluid models. Problem-specific regularization adds data knowledge such as segmentation maps, organ boundaries, or clinical context, making the constraint spatially adaptive; this family covers multi-structure registration, organs with sliding or cyclic motion, and images with topological change. Learned regularization parameterizes the regularizer itself with a machine or deep learning model, either learning local smoothing and discontinuities, learning low-dimensional feasible deformation spaces through PCA models or autoencoders, or learning test-time adaptive regularization weights. The review further argues that prior information increases from family I to III, that most problem-specific and learned methods extend model-based ones, and that measured against this map the field shows a strong default to global $L^2$-norm smoothing — 14 of 21 Learn2Reg 2022 methods — while sliding motion, local rigidity, cyclic motion, and physics-inspired regularization remain underrepresented in learning-based frameworks.","pith_inferences":["Should the taxonomy become the field's standard reference, new hybrid regularizers that combine learned components with model-based constraints — a direction the review itself anticipates — would press hardest on the boundary between families II and III, which is where most future growth is likely to occur.","The review notes that learned deformation spaces inherit the properties of whatever algorithm generated their training deformations; an extension of that observation is a circularity risk, namely that data-driven regularization may silently re-encode the global-smoothness bias it is meant to escape, and auditing learned regularizers for behavior beyond their training-generating regularization woul","The taxonomy could be operationalized as a decision rule — no extra data leads to model-based choices, segmentations or physiological context to problem-specific ones, large deformation datasets to learned ones — and a benchmark comparing these routes on sliding-motion and topology-change tasks would turn the review's qualitative claims into quantitative ones.","Extending the 14-of-21 Learn2Reg count into a broader census of recent registration papers would test whether the identified overreliance on plain smoothing persists as the field evolves."],"forward_implications":["A researcher facing a new registration task can locate the relevant regularization family and transfer solutions from analogous applications instead of defaulting to the standard smoothing term.","The documented default to global $L^2$-norm smoothing — 14 of 21 methods in the Learn2Reg 2022 challenge — implies that leading learning-based frameworks are likely leaving anatomical plausibility unaddressed, particularly for sliding motion and locally rigid structures.","Sliding-motion, local-rigidity, cyclic-motion, incompressibility, and physics-inspired regularization are the least transferred into learning-based registration, making them concrete targets for new methodological work.","Test-time regularization — tuning the regularization weight at inference through hypernetworks or conditional normalization layers — is presented as the bridge between instance-specific conventional tuning and fast learning-based inference.","Evaluation based on the fraction of negative Jacobian determinants can mislead, since anatomically realistic sliding motion can increase folding; the review argues for targeted measures such as maximum shear and deformation-space reconstruction error."],"supporting_citations":[{"why":"Defines regularization as one of the fundamental building blocks of registration and supplies the objective function and notation around which the whole review is organized.","marker":"(Rueckert and Schnabel, 2010)"},{"why":"Origin of the diffusion-type $L^2$-norm smoothness penalty that the review identifies as the default regularization in modern frameworks.","marker":"(Horn and Schunck, 1981)"},{"why":"VoxelMorph, the canonical learning-based registration framework whose reliance on global smoothing exemplifies the default approach the review critiques.","marker":"(Balakrishnan et al., 2019)"},{"why":"Learn2Reg 2022 challenge, the source of the statistic that 14 of 21 top methods use $L^2$-norm smoothing, grounding the overreliance claim.","marker":"(Hering et al., 2023)"},{"why":"General registration survey cited as evidence that existing reviews cover regularization only briefly, justifying the claimed gap.","marker":"(Viergever et al., 2016)"},{"why":"Deep-learning registration survey against which the review positions its taxonomy as the first structured treatment of regularization.","marker":"(Haskins et al., 2020)"},{"why":"Stationary velocity field and log-Euclidean diffeomorphic framework, the model-based method most widely transferred into learning-based registration.","marker":"(Arsigny et al., 2006)"},{"why":"Free-form deformation with B-splines, the classic implicit smoothness-by-parameterization method used across both conventional and learning-based registration.","marker":"(Rueckert, 1999)"},{"why":"HyperMorph, one of the two foundational methods defining the learned test-time regularization subcategory.","marker":"(Hoopes et al., 2021)"},{"why":"Conditional instance normalization approach that, with HyperMorph, establishes learned test-time regularization as an emerging approach.","marker":"(Mok and Chung, 2021a)"}],"fun_headline_variants":["Three-family taxonomy maps all registration regularizers","Registration's dirty secret: most methods stick to basic smoothing","From handcrafted to learned: a new map for regularization","Review reveals three pillars of image registration regularization"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The value of the whole map rests on the literature search being essentially complete and on every regularization method fitting cleanly into exactly one of the three families, even though the search itself is described only briefly and unlisted.","fun_headline_variants_meta":{"raw":{"variants":["Three-family taxonomy maps all registration regularizers","Registration's dirty secret: most methods stick to basic smoothing","From handcrafted to learned: a new map for regularization","Review reveals three pillars of image registration regularization"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000219,"raw_usage":{"total_tokens":1488,"prompt_tokens":1036,"completion_tokens":452,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":652,"completion_tokens_details":{"reasoning_tokens":390}},"tokens_in":652,"tokens_out":452,"duration_ms":4571,"temperature":1.0,"reasoning_tokens":390,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T11:07:39.324570+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the review's own search protocol independently — the same three databases, the same keyword combinations, the same October 2024 cutoff — and compare the recovered method set against the review's tables; if a substantial body of published regularization methods is missing or cannot be assigned to exactly one of the three families, the claim of comprehensive structured coverage fails. A cheap first probe is to count how many methods cited within the review's own reference list never appear in its category tables.","supporting_citations":[],"review_version":1}