{"id":"57b9ee6b-60cc-43b9-8a9a-4812892392ce","arxiv_id":"1908.02738","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A learning framework jointly estimates deformable templates and diffeomorphic registration networks, enabling rapid synthesis of conditional templates such as brain atlases for a given age.","lead":"Researchers trained a neural network to jointly learn a deformable image template and fast image-to-template alignment, and to generate templates conditioned on attributes such as age or class. The approach could make custom medical atlases cheap to build, replacing costly iterative algorithms that take days or weeks.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Conditional templates are never tested against a non-circular outcome: the neuroimaging Dice comparison is self-admittedly non-comparable and shows no gain over the unconditional template, so the central claim of useful age-conditioned atlases is currently unsupported.","rationale":"The reader's weakest_assumption (single diffeomorphic template) is a genuine modeling limitation, but it is standard for deformable-template methods and the datasets used here do not clearly violate it. I therefore treat it as secondary. The more load-bearing gap is that the paper's central novelty—conditional templates—has no non-circular quantitative validation. The toy experiments with controlled scale and rotation on MNIST/QuickDraw are encouraging and do demonstrate the learning mechanism in a synthetic setting, which is independent evidence that the framework can work. But the transition to the neuroimaging application, which motivates the abstract, rests on visual inspection and self-referential trends. The centrality metrics in Figure 6 also coincide with terms in the training loss, so they do not independently establish that the templates are anatomically central. This is an evaluation gap rather than a mathematical inconsistency, so I do not recommend rejection; the method may well be correct. However, the acceptance conditions should include the matched-vs-mismatched conditional template test above. This aligns only partially with the reader's weakest_assumption: we agree that the current evidence is insufficient, but the precise soft spot I identify is the absence of a functional test of conditioning, not the diffeomorphic model assumption.","tokens_in":13784,"tokens_out":11503,"duration_ms":142550,"concrete_test":"On the held-out 100 test subjects with FreeSurfer segmentations, register each subject to (a) its age/sex-matched conditional template, (b) an age-mismatched conditional template (e.g., the template at age ±20 years), and (c) the unconditional template, using the same trained registration network. Compute Dice over the 30 FreeSurfer labels and compare ventricle and hippocampus volumes estimated from warped template segmentations against the subject's native FreeSurfer volumes. If matched conditional templates do not improve Dice or reduce volume-estimation bias relative to mismatched or unconditional templates, the conditional component is not functionally validated and the central claim is weakened.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's main novelty is the conditional template function. Its value is supported by three pieces of evidence, none of which is decisive. (1) Figure 13 reports ventricle/hippocampus volume trends computed by warping training segmentations into the conditional templates; this is a description of the model's own outputs, not an independent validation that conditioned templates improve alignment or analysis. (2) The only quantitative neuroimaging comparison (Section 4.2) is flagged by the authors themselves: 'these numbers may not be directly compared' because the baseline atlas used a different external dataset and labeling pipeline, so the abstract-level claim of 'atlases similar in quality and utility to a widely used atlas' is unsupported. (3) The conditional template's Dice (0.795) is not better than the unconditional template's (0.800), so there is no measured benefit from conditioning. The loss in Eq. 6 contains no term that forces attribute-related geometry into the template rather than into per-image deformations; the deformation-magnitude and smoothness priors are the only incentive. Absent a head-to-head matched-vs-mismatched conditional template test on held-out data, the visual age trends could be produced by a model that leaves attribute information in the registration fields while the template network emits a smooth compromise.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a probabilistic framework for learning deformable templates from image collections, in which each image is modeled as a diffeomorphic deformation of a template that may itself be a function of observed attributes. The authors derive a maximum-likelihood objective (Eq. 6) combining an image likelihood, a deformation-magnitude prior, and a smoothness prior, and optimize it end-to-end with a template-generation network and a registration network. They present experiments on MNIST and QuickDraw with controlled scaling/rotation attributes, and on a large multi-site brain MRI dataset, where they build unconditional and age/sex-conditioned templates. The paper claims that the learned templates are comparable in quality to a widely used atlas and that conditional templates capture age-related anatomical changes such as ventricle growth and hippocampal shrinkage.","tokens_in":14120,"tokens_out":4604,"duration_ms":45573,"significance":"The maximum-likelihood derivation is clear, the method is general, and the implementation (as part of VoxelMorph) with code and atlases released is a strength. The conditional-template formulation is a useful contribution: it provides a single network that can generate on-demand templates for arbitrary observed attributes without subdividing the population, and the MNIST experiments demonstrate that the model can learn attribute-driven geometric variability and generalize to sparse or held-out attribute values. The neuroimaging results, however, do not yet substantiate the claim that conditional templates improve or match existing atlases in a quantitatively meaningful way, because the only quantitative comparison is explicitly not apples-to-apples and shows no gain over the unconditional template. The value of the contribution therefore rests mainly on the toy experiments and visual/descriptive neuroimaging evidence.","major_comments":[{"comment":"The Dice comparison cannot support the abstract-level claim that the learned atlases are 'similar in quality and utility to a widely used atlas.' The authors state that 'these numbers may not be directly compared' because the baseline atlas and segmentations were produced with an external dataset and a different labeling pipeline, while their template labels come from FreeSurfer on their own training images. Moreover, the conditional template achieves Dice 0.795 ± 0.116, numerically below the unconditional template's 0.800 ± 0.110, with no significance test. A matched comparison on the same test subjects, using the same segmentation protocol for all templates, and a conditional-versus-unconditional comparison, is required to support the utility claim.","section":"Section 4.2, Dice evaluation"},{"comment":"The age-related volume trends in Figure 13 are computed by warping training segmentations into the model's own conditional templates and measuring volumes of those templates. Because the template network and the deformation network are optimized jointly and the loss in Eq. (6) contains no term that explicitly forces attribute-related geometry into the template, these trends are a description of the learned model's outputs rather than an independent validation that conditional templates capture true anatomical variability. A non-circular test could compare registration accuracy or segmentation accuracy on held-out images when using matched versus mismatched conditional templates (e.g., age-appropriate versus age-inappropriate templates).","section":"Section 4.2, Fig. 13"},{"comment":"The centrality metric (mean displacement norm) is directly penalized in the training objective through the term -γ‖ū‖² in Eq. (6), so reporting lower centrality for the proposed method as a success criterion is partly circular. The comparison with the decoder baseline is informative about the objective being optimized, but it does not establish that the templates are better in an independent sense. The more decisive quantitative evidence in this section is the MSE and Jacobian behavior, which should be emphasized, and centrality should be presented as an objective-matching check rather than an external quality measure.","section":"Section 4.1.1, centrality metrics"},{"comment":"The generative model assumes that each image is a spatially deformed version of a single conditional template under a diffeomorphic stationary-velocity-field deformation. In heterogeneous clinical cohorts containing lesions, tumors, or large contrast differences, this assumption is violated and the learned template will be a compromise that cannot be registered accurately to all images. Since the paper motivates clinical applications, the authors should either add a stress test with such images (e.g., synthetic lesions or outlier scans) or explicitly scope the claims to populations that satisfy the mutual-diffeomorphism assumption.","section":"Section 3.1, generative assumption"}],"minor_comments":[{"comment":"The running-average approximation of ū is introduced only in the text; the notation in Eq. (6) treats ū as if it were the full dataset mean. The approximation should be made explicit in the equation or immediately after it.","section":"Section 3.2, Eq. (6)"},{"comment":"The text says the conditional model was trained 'using only the ADNI and ABIDE datasets' but the overall dataset includes many sources; clarify which splits were used for the unconditional and conditional models and how the 250 test subjects were selected.","section":"Section 4.2"},{"comment":"No error bars or confidence intervals are given for the volume-age trends, and the y-axis units ('x1000 voxels') are potentially confusing; specify whether these are raw voxel counts and report variability across subjects or templates.","section":"Section 4.2, Figure 13"},{"comment":"The term 'unbiased population templates' is used for a template whose mean deformation is small; this is not the same as statistical unbiasedness of an estimator. Consider rewording to 'central' templates.","section":"Sections 2.2 and 3.1"},{"comment":"There are several typos and notation issues: 'levarges' should be 'leverages' in Section 2.2, and in Section 4.1.1 the expression '|Jφ(p)|≤0' should be clarified to mean Jacobian determinants at or below zero, with the intended interpretation of non-topology-preserving pixels.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"I recommend major revision. The method is sound in principle and the toy experiments are encouraging, but the neuroimaging evaluation must be strengthened to support the central claim that conditional templates are useful for clinical applications. If the authors provide matched conditional-versus-unconditional comparisons and independent validation, the paper could become acceptable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Core contribution is real: after a day of training, a single network can synthesize a deformable template conditioned on any attribute vector in under a second, while also giving a diffeomorphic registration to that template. The probabilistic derivation (Eqs 1-6) is clean, and the MNIST/QuickDraw experiments show the conditional function actually picks up simulated scale and rotation. Code and atlases are public, which is a plus.\n\nThe soft spot is the neuroimaging evaluation. The only quantitative comparison, Dice in Sec 4.2, is flagged by the authors as not directly comparable because the baseline atlas came from a different pipeline and dataset. The conditional template scores 0.795, the unconditional 0.800; there is no significance testing and no matched comparison of conditional versus unconditional on held-out data. Figure 13's age-volume trends are model outputs, not independent validation. So the abstract/conclusion claim that this produces atlases \"similar in quality and utility\" to a widely used atlas is not supported by the evidence. The MNIST analyses also use centrality metrics that appear in the loss, though the baselines make those comparisons useful.\n\nWhat is genuinely new: jointly learning the template and the registration network, and a conditional template function over continuous or discrete attributes. Prior work optimized atlases iteratively or registered to a fixed template. The missing-attribute and latent-attribute experiments are thoughtful and show the model can generalize. The neuroimaging weakness is a gap in evaluation, not a flaw in the method.\n\nThe main thing missing is one clean head-to-head: same data, same pipeline, compare (a) unconditional template, (b) age-matched conditional template, and (c) age-mismatched conditional template, and report Dice and deformation norms with error bars. That would either validate or kill the clinical-utility claim.\n\nI would send this to reviewers. It is a genuine contribution and the toy evidence is strong, but it needs revision to bring the neuroimaging claims in line with what was measured. The reader's \"conditional\" verdict is right; the stress-test note is a fair description of the current evidentiary state, not an indictment of the method.","headline":"A clean, novel method for learning conditional templates jointly with a registration network; the MNIST evidence is solid, but the neuroimaging evaluation does not yet show that conditioning helps.","tokens_in":14582,"tokens_out":3137,"would_cite":true,"duration_ms":38166,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that a single convolutional network can jointly estimate a deformable template for a whole population—or one conditioned on attributes such as age and sex—and provide fast diffeomorphic alignment to any new image.","keywords":["deformable templates","conditional atlases","diffeomorphic image registration","probabilistic models","neuroimaging","convolutional networks","template estimation","brain MRI"],"falsifier":"Train the same model on a cohort that includes focal lesions or large scanner contrast differences and compare per-subject alignment quality, for example Dice overlap or reconstruction error, against a model trained only on homogeneous controls. If the conditional template cannot be deformed to match the atypical images with topology-preserving smooth fields, and those subjects' metrics fall well below the control average, the central modeling assumption is violated.","tokens_in":13621,"feed_emoji":"🧠","tokens_out":6629,"duration_ms":69118,"temperature":0.7,"pith_summary":"The paper introduces a learning-based alternative to traditional deformable-template construction, which iteratively registers an entire population to a central image and can take days to weeks. It proposes a probabilistic generative model in which every image is a spatially deformed version of a template, and a neural network that, in one pass, outputs a template conditioned on given attributes and the velocity field that aligns the template to the input image. If the method works as claimed, on-demand templates can be synthesized in under a second, and the same network returns both the template and the deformation to any new image. The paper demonstrates this on handwritten-digit and sketch datasets and on a large brain MRI collection, showing unconditional and age-conditioned templates with anatomical trends such as ventricle growth and hippocampal shrinkage.","feed_headline":"One network builds age-dependent brain templates in seconds","feed_subtitle":"A probabilistic model learns an on-demand template, tuned by attributes, plus fast diffeomorphic alignment.","key_machinery":"The central objects are a conditional template network $g_{t,\\theta_t}(a_i) = t$, implemented as a decoder, and a registration network $g_{v,\\theta_v}(t, x_i) = v$, implemented as a U-Net, joined by the stationary-velocity-field parametrization of diffeomorphisms: integrating $v$ via scaling and squaring gives the deformation $\\phi_v$, and warping the template by $\\phi_v$ reconstructs the image. The generative model ties these together through the likelihood $p(x_i \\mid v_i; a_i) = \\mathcal{N}(x_i;\\, f_{\\theta_t}(a_i) \\circ \\phi_{v_i},\\, \\sigma^2 I)$ and the velocity prior $p(V) \\propto \\exp(-\\gamma\\|\\bar{u}\\|^2)\\prod_i \\mathcal{N}(u_i;\\, 0,\\, \\Sigma_u)$, where the mean-displacement term enforces an unbiased central template and the Laplacian-based covariance encourages smooth deformations. Minimizing the negative log likelihood with stochastic gradients updates template and registration networks jointly, sidestepping expensive iterative pairwise registration.","core_discovery":"The central claim is that template estimation and image alignment can be cast as a single maximum-likelihood learning problem. For each image $x_i$, the model posits that $x_i$ is generated by warping a conditional template $t = f_{\\theta_t}(a_i)$ through a diffeomorphism $\\phi_{v_i}$ parameterized by a stationary velocity field $v_i$, with an additive Gaussian or normalized-cross-correlation likelihood and a deformation prior that penalizes the mean displacement and encourages smoothness. A network $g_\\theta(x_i, a_i) = (v_i, t)$ is trained end-to-end by minimizing the negative log likelihood, learning at once a template function and a fast registration network; at test time, both the template and the deformation are obtained in a forward pass, and inverse deformations come from integrating the negative velocity field. The experiments claim that this produces central templates requiring smaller deformations than instance-based or decoder-only baselines, and conditional brain templates consistent with known age-related anatomical changes.","pith_inferences":["Beyond the paper, the learned template function could be inverted to estimate attributes such as age from a scan, by finding the attribute value whose template best aligns with the input; the paper mentions this only as future work.","The latent-attribute experiment suggests a fully unsupervised version of this framework could build class- or mode-conditional templates without observed labels, effectively discovering geometric factors of variation in a dataset.","A testable extension would apply the model to multi-site clinical data with strong scanner or contrast differences; under the current likelihood, such appearance variation might be absorbed into the template and deformations rather than being modeled as noise.","If conditional template functions are smooth in the attributes, the framework implicitly yields a generative model of anatomy along those axes, so one could synthesize new population samples by deforming a conditional template with sampled velocity fields."],"forward_implications":["A single trained model provides both a template and a deformation field for any new image in one forward pass, making on-demand atlas construction practical in clinical settings where no pre-existing template is available.","Conditional templates let the same data support many subpopulation atlases, such as an age- and sex-specific brain template, without subdividing the dataset or arbitrarily thresholding continuous attributes.","Because the template is a learned function of attributes, the model can synthesize templates for attribute values that were sparsely observed or held out during training, interpolating age-related anatomy across the whole range.","Learning to represent images up to a deformation means the method captures geometric variability aligned with the conditioning attributes, and can reduce confounding effects when those attributes are supplied."],"supporting_citations":[{"why":"Supplies the stationary-velocity-field diffeomorphic registration framework and the negative-velocity inverse used for deformations.","marker":"[5]"},{"why":"Provides the U-Net registration architecture that the network $g_v$ is based on.","marker":"[9]"},{"why":"Introduces the unsupervised probabilistic diffeomorphic registration approach with scaling-and-squaring integration and the Laplacian deformation prior reused here.","marker":"[17]"},{"why":"Gives the unsupervised registration baseline and the online atlas with segmentations used for Dice comparison in the neuroimaging experiment.","marker":"[8]"},{"why":"Defines the unbiased diffeomorphic atlas construction problem that this learning-based method targets.","marker":"[40]"},{"why":"Provides the statistical dense deformable template estimation framework that this probabilistic model extends.","marker":"[3]"},{"why":"Assembles the multi-site brain MRI dataset used in the neuroimaging experiments.","marker":"[20]"}],"fun_headline_variants":["Conditional deformable templates from a single network","One network learns templates and aligns in one forward pass","Fast conditional brain atlas and warping via deep learning","Age-conditional templates and diffeomorphic alignment with CNNs"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing assumption is that every image in the dataset can be represented as a single smoothly deformed version of a shared conditional template; images that cannot be diffeomorphically matched to one central shape, such as brains with tumors, lesions, or strong scanner-related contrast differences, would force the template into a compromise and bias the deformation estimates.","fun_headline_variants_meta":{"raw":{"variants":["Conditional deformable templates from a single network","One network learns templates and aligns in one forward pass","Fast conditional brain atlas and warping via deep learning","Age-conditional templates and diffeomorphic alignment with CNNs"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000763,"raw_usage":{"total_tokens":3382,"prompt_tokens":936,"completion_tokens":2446,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":552,"completion_tokens_details":{"reasoning_tokens":2382}},"tokens_in":552,"tokens_out":2446,"duration_ms":16933,"temperature":1.0,"reasoning_tokens":2382,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T14:36:21.553161+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train the same model on a cohort that includes focal lesions or large scanner contrast differences and compare per-subject alignment quality, for example Dice overlap or reconstruction error, against a model trained only on homogeneous controls. If the conditional template cannot be deformed to match the atypical images with topology-preserving smooth fields, and those subjects' metrics fall well below the control average, the central modeling assumption is violated.","supporting_citations":[{"cited_title":"Ashburner","cited_arxiv_id":null,"evidence_quote":"Supplies the stationary-velocity-field diffeomorphic registration framework and the negative-velocity inverse used for deformations."},{"cited_title":"Balakrishnan, A","cited_arxiv_id":null,"evidence_quote":"Provides the U-Net registration architecture that the network $g_v$ is based on."},{"cited_title":"Dalca, G","cited_arxiv_id":null,"evidence_quote":"Introduces the unsupervised probabilistic diffeomorphic registration approach with scaling-and-squaring integration and the Laplacian deformation prior reused here."},{"cited_title":"Balakrishnan, A","cited_arxiv_id":null,"evidence_quote":"Gives the unsupervised registration baseline and the online atlas with segmentations used for Dice comparison in the neuroimaging experiment."},{"cited_title":"Unbiased diffeomorphic atlas construction for computational anatomy","cited_arxiv_id":null,"evidence_quote":"Defines the unbiased diffeomorphic atlas construction problem that this learning-based method targets."},{"cited_title":"Towards a coherent statistical framework for dense deformable template estimation","cited_arxiv_id":null,"evidence_quote":"Provides the statistical dense deformable template estimation framework that this probabilistic model extends."},{"cited_title":"Dalca, J","cited_arxiv_id":null,"evidence_quote":"Assembles the multi-site brain MRI dataset used in the neuroimaging experiments."}],"review_version":1}