{"id":"44428115-d9a7-455b-a3d2-b3d378e49dd8","arxiv_id":"2608.09460","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":8,"one_line_summary":"CAN-FLOW, a two-step conditional normalizing flow generator trained on LDDMM momenta from 2,208 UK Biobank hearts, produces sex-, age-, and BMI-conditioned biventricular anatomies whose variability matches the real cohort more closely than cVAE baselines.","lead":"CAN-FLOW is a new two-step generative model that creates realistic virtual heart anatomies conditioned on sex, age, and body mass index, trained on 2,208 healthy UK Biobank participants. It addresses a bottleneck in cardiac digital twin research: building virtual patient cohorts that preserve real anatomical variability without sharing individual medical images.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Most distributional-fidelity results compare synthetics against the full real cohort, including the 70% training split; the claimed advantage may reflect memorization rather than held-out generalization.","rationale":"The reader's weakest assumption identifies the same load-bearing concern: the evaluation protocol does not demonstrate that cohort-level metrics are restricted to held-out subjects. The full text confirms this: Section 4.1 describes the split, but Sections 2.2-2.6 and Appendix A.5 repeatedly refer to 'the real cohort' without specifying that it is the test split. This matters because the paper's headline claim is distributional fidelity of generated cohorts, not reconstruction accuracy on training data. If the reference distribution includes training anatomies, the metrics can be artificially favorable for any sufficiently flexible model, and the comparison to cVAEs becomes a test of overfitting rather than generalization. The proposed check directly settles this by rerunning the main evaluations on the test split only. The reader's verdict of conditional acceptance remains appropriate: the concern is concrete and addressable, but it is not fatal to the method's plausibility, and the paper has independent support from ablations over flow hyperparameters (Appendix B.4) and latent dimensionality (Appendix B.5). I do not see a stronger load-bearing threat than the train/test overlap, and I agree with the reader's assessment.","tokens_in":25768,"tokens_out":2384,"duration_ms":26586,"concrete_test":"Recompute Figures 4, 6, 8, and 9 (and the Appendix B tables) using only the held-out 15% test split from Section 4.1 as the real reference cohort, with synthetic cohorts matched to the test-set size and metadata distribution; if CAN-FLOW's advantage over cVAEs in KL/WD, MMD/coverage, and PCA ratios disappears or reverses on held-out subjects, the distributional-fidelity claim is not established. The paper should also explicitly state whether any training subjects are included in the reference distributions for all cohort-level metrics.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that CAN-FLOW better reproduces the real population distribution than cVAE baselines. That claim requires comparing synthetic cohorts to a reference distribution that was not used to train the model. Section 4.1 defines a 70/15/15 train/validation/test split, but the evaluation sections (2.2 through 2.6, and Appendix A.5) never state that the real reference cohort is restricted to the held-out test split. Appendix A.5 says 'Age and BMI were sampled uniformly from the ranges observed in the real cohort' and describes subsampling '600 anatomies from the real cohort' (A.5.2, A.5.4), without mentioning exclusion of training subjects. If the reference distribution includes the training anatomies, then the reported KL divergences, Wasserstein distances, MMD/coverage scores, and PCA-based variability ratios partly measure how well the model reproduces data it was explicitly trained on. A flexible conditional normalizing flow with 15 Glow blocks and a metadata-dependent prior can overfit the training set, so low distances to the training cohort would not be evidence of generalization to new subjects. This is not an internal inconsistency, but it is a direct validity threat to the paper's headline comparison against cVAEs. A secondary concern is that the 66 Mahalanobis outliers removed in Appendix A.2 may include legitimate tail anatomies; because the same trimmed cohort serves as both training and reference, the evaluation cannot measure preservation of the tails the paper claims to preserve.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces CAN-FLOW, a two-stage conditional generative model for biventricular cardiac anatomy. In the first stage, an autoencoder compresses LDDMM momenta into a geometry-only latent space. In the second stage, a conditional normalizing flow with a metadata-dependent prior, affine injectors, and conditional coupling layers models the distribution of these latents given sex, age, and BMI. The authors compare CAN-FLOW against cVAEs trained over a range of β values on a healthy UK Biobank cohort of 2,208 subjects. Evaluation covers visual plausibility, clinical phenotype distributions (KL divergence, Wasserstein distance), metadata-stratified phenotype trends, subgroup variability, PCA-based shape variability in momenta space, and point-cloud MMD/coverage. The paper reports that CAN-FLOW outperforms cVAEs on most distributional fidelity metrics, with ablations in Appendix B.4 and B.5 showing robustness to architectural choices.","tokens_in":26119,"tokens_out":5747,"duration_ms":56553,"significance":"If the distributional fidelity results survive a proper held-out evaluation, CAN-FLOW is a valuable contribution to virtual cohort generation for cardiac digital twins. The two-step design—decoupling representation learning from conditional density estimation—is a conceptually clean alternative to cVAEs, and the paper evaluates it with a broad set of complementary metrics (KL, Wasserstein, MMD, coverage, PCA ratios). The ablation studies in Appendix B.4 and B.5 are a genuine strength, as they show the advantage is not tied to a single fine-tuned configuration. The paper also clearly defines a train/validation/test split in Section 4.1, indicating awareness of generalization. The central weakness is that the reported quantitative comparisons appear to use the full post-outlier cohort, including the training split, as the real reference distribution, which undermines the generalization claim until re-run on held-out data.","major_comments":[{"comment":"The distributional evaluation uses the full post-outlier cohort as the real reference, not the held-out test split. Section 4.1 defines a 70/15/15 train/validation/test split, but Sections 2.2 through 2.6 and Appendix A.5 never state that the real cohort subsets are restricted to the test set—for example, the 600 subsampled anatomies in Figure 9 and Appendix A.5.2, the two PCA subsets in Appendix A.5.4, and the subgroup in Section 2.2. Since a flexible normalizing flow with 15 Glow blocks can overfit the training data, the reported KL divergences, Wasserstein distances, MMD/coverage scores, and PCA variability ratios likely reflect, at least in part, memorization of training anatomies rather than generalization to new subjects. The authors should re-run all cohort-level metrics against the held-out test subjects, or at minimum report the train and test metrics separately, and show that CAN-FLOW's advantage over cVAEs persists. This is load-bearing for the paper's headline claim that CAN-FLOW better reproduces the real population distribution.","section":"Section 4.1 and Appendix A.5"},{"comment":"The paper motivates preserving rare but plausible tail anatomies (Introduction, Section 2.3), yet it removes 66 Mahalanobis outliers before training (Appendix A.2) and then uses the same trimmed cohort as the reference in every evaluation. Consequently, the reported metrics only measure fidelity to the post-trimming distribution; they cannot validate preservation of the original extreme tails. The authors should explicitly state whether the removed outliers are included in any reference distribution, and either include them in an additional tail-focused analysis or temper the tail-preservation claim accordingly.","section":"Appendix A.2 and Discussion (tail preservation)"},{"comment":"The overall-population comparisons sample age and BMI uniformly from the ranges observed in the real cohort, separately by sex, while the real cohort reference retains its natural joint metadata distribution. This means the synthetic and real cohorts have different metadata marginals, so the reported KL and Wasserstein distances for the whole cohort confound anatomical fidelity with metadata-marginal mismatch. The authors should either sample the synthetic metadata from the real joint metadata distribution (for example, by bootstrapping real metadata vectors) or restrict the overall distributional claims to the metadata-stratified analyses, which are less sensitive to this issue.","section":"Appendix A.5 and Sections 2.3, 2.6"}],"minor_comments":[{"comment":"The notation for the composition of flow blocks is typeset awkwardly; consider writing f_phi = f_phi^{(N_b)} ∘ ... ∘ f_phi^{(1)} to make the composition order explicit.","section":"Section 4.3, Eq. (1)"},{"comment":"Appendix A.1 defines N_p = 2274 as the number of anatomies, while Section 4.1 reports 2208 after outlier removal; please clarify whether N_p refers to the pre- or post-outlier count and make the numbering consistent throughout.","section":"Appendix A.1 vs Section 4.1"},{"comment":"The subgroup 'males older than 58 years' is described as containing 674 subjects; please state whether this is the full cohort count or a split-specific count, and report the synthetic cohort size used for comparison.","section":"Section 2.2, Figure 3"},{"comment":"The phrase 'the largest value in each column is marked in blue' is confusing because lower KL and Wasserstein values are better; presumably the marking indicates the worst value. Please reword to avoid ambiguity.","section":"Appendix B.4, Tables B.2–B.4"},{"comment":"Minor typographical and formatting issues include inconsistent number formatting (e.g., '2,208' in the abstract vs '2208' in Section 4.1) and several superscript/subscript renderings in Section 4.3 that should be cleaned up.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The main validity threat is the evaluation-on-training-data issue, which is fixable by re-running the distributional metrics on the held-out test split. If the authors can show that CAN-FLOW's advantage persists on held-out data, the paper would be a solid contribution. The metadata-marginal confound is also worth addressing, though it is less severe. The paper does not currently provide code or a full data-access statement, but that is promised for publication and is not a blocker at this stage."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The two-step split is the real contribution here: a geometry-only autoencoder on LDDMM momenta followed by a conditional normalizing flow on the latent codes. That cleanly separates representation learning from density estimation, and it directly addresses the known cVAE failure mode of compressing subgroup variance under a shared prior. The paper is also honest and well engineered; the ablations in Appendix B.4 and B.5 show the advantage persists across activation functions, block counts, embedding sizes, and latent dimensions, which is more than most papers in this space provide.\n\nThe main soft spot is exactly what the stress-test flags. The paper defines a 70/15/15 split in Section 4.1, but the distributional evaluations in Sections 2.2 through 2.6 and Appendix A.5 compare synthetic cohorts to subsamples of the real cohort without ever stating that training subjects are excluded. If the reference distribution includes the 70% training split, then the KL divergences, Wasserstein distances, MMD, and coverage partly measure how well the flow reproduces data it was trained on. A flexible 15-block Glow with a metadata-dependent prior can overfit, so the reported advantage over cVAEs could be inflated. This is not an internal contradiction, and it is fixable by rerunning the metrics on the held-out test split. But without that rerun, the headline claim of \"better reproduces the real population distribution\" is not fully supported.\n\nSecondary issues are proportionate: the radar plots have no error bars or significance tests, though the PCA ratios in Figure 8 do report ±σ over five splits; the 66 Mahalanobis outliers may include legitimate tails, and the same trimmed cohort is used as both training and reference; and the closest prior conditional-flow baseline (Dou et al. [35]) is not benchmarked, though the cVAE sweep over β is a fair baseline set. None of these are fatal. The paper is clear about its scope — healthy UK Biobank, end-diastolic only — and the discussion of limitations is honest.\n\nThis paper deserves a serious referee. The method is new and plausible, the evidence is consistent, and the evaluation gap is addressable rather than structural. If the authors restrict the reference cohort to the held-out split, add uncertainty estimates, and ideally benchmark against a conditional flow baseline, the result would be a solid contribution to virtual cohort generation. I would bring it to a reading group and, after the revision, cite it in my own work.","headline":"The two-step split is a real contribution and the paper is well engineered, but the distributional claims need a held-out reference cohort before they fully land.","tokens_in":26657,"tokens_out":2647,"would_cite":true,"duration_ms":27454,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"CAN-FLOW, a two-step normalizing-flow framework, generates metadata-conditioned biventricular anatomies whose distributional spread matches a real healthy cohort more closely than conditional variational autoencoders.","keywords":["synthetic clinical data","cardiac anatomy generation","virtual cohorts","cardiac digital twins","conditional generative models","normalizing flows","biventricular anatomy","UK Biobank"],"falsifier":"Recompute the KL divergences, Wasserstein distances, point-cloud coverage, and PCA variability ratios using only the 15% held-out test subjects as the real reference population. If CAN-FLOW's advantage over the cVAE baselines shrinks or reverses on out-of-training anatomies, the paper's central distributional-fidelity claim is refuted.","tokens_in":25576,"feed_emoji":"🫀","tokens_out":7021,"duration_ms":64157,"temperature":0.7,"pith_summary":"CAN-FLOW is a generative model for biventricular heart anatomy that tries to solve a specific problem: virtual cohorts for cardiac digital twins and in silico trials need to reproduce the full, metadata-dependent spread of real anatomies, not just average shapes. The paper claims that separating geometry-only representation learning from conditional density modeling achieves this. An autoencoder compresses diffeomorphic shape momenta into an unregularized latent space, and a conditional normalizing flow then models how that latent space depends on sex, age, and BMI. Across clinical phenotypes, subgroup-stratified distributions, point-cloud coverage, and high-dimensional shape variability, CAN-FLOW matched a healthy UK Biobank cohort better than cVAE baselines, which produced more homogeneous cohorts and underrepresented distribution tails. If correct, this makes synthetic but realistic cardiac anatomy cohorts feasible without sharing individual imaging data.","feed_headline":"CAN-FLOW reproduces heart anatomy variability VAEs miss","feed_subtitle":"Two-step flow model generates virtual heart cohorts that match real phenotype distributions better than cVAE baselines.","key_machinery":"The load-bearing mechanism is the two-step decoupling of representation learning from conditional density estimation. Anatomies are encoded as LDDMM initial momenta $\\mu_0 \\in \\mathbb{R}^{3 \\times 720}$; an unregularized autoencoder compresses these momenta into a 44-dimensional latent $z$, and a Glow-style conditional normalizing flow $f_\\phi(z,c)$ learns the metadata-dependent density of $z$, with a learnable conditional base $\\mathcal{N}(\\mu_\\phi(c), \\Sigma_\\phi(c))$, affine injectors, and conditional affine coupling layers. At generation time the flow samples from the conditional base, the inverse flow maps it to $\\tilde{z}$, and geodesic shooting converts decoded momenta back into surface meshes. This design is what lets the model vary the latent distribution with metadata without forcing a shared prior during representation learning.","core_discovery":"The paper's central claim is that the reason previous conditional anatomy generators underrepresent population variability is architectural: conditional variational autoencoders tie representation learning to a fixed, metadata-agnostic Gaussian prior, so the latent space is regularized toward a single shared distribution. CAN-FLOW breaks that coupling. It first learns an unconstrained, geometry-only latent representation of diffeomorphic shape momenta, then fits a conditional normalizing flow that maps sex, age, and BMI to a learnable Gaussian base distribution and an invertible transformation of that base. On a healthy cohort of 2,208 UK Biobank biventricular meshes, the paper reports that CAN-FLOW outperformed the strongest cVAE baselines on KL divergence for all clinical phenotypes, on most Wasserstein-distance comparisons, on sex- and age-conditioned phenotype distributions, on point-cloud coverage, on within-subgroup spatial variability, and on PCA-based shape-momenta variability, while cVAEs achieved lower minimum matching distance but produced visibly more homogeneous cohorts.","pith_inferences":["Editorial inference: Because the paper's evaluation uses the same post-outlier cohort for training and reference, the distributional metrics may partly reflect memorization; a held-out-only re-analysis is the natural check.","Editorial inference: Removing 66 Mahalanobis outliers before training likely deletes the most extreme legitimate morphologies, so the claim of preserved tails applies to the retained distribution, not to the rarest real anatomies.","Editorial inference: A stronger test of generalizability would train on one imaging site or population and evaluate on an external cohort; otherwise the learned healthy-anatomy distribution remains UK Biobank-specific.","Editorial inference: The static end-diastolic scope means CAN-FLOW does not yet address motion or phase-consistent deformation, so extending it to four-chamber or time-resolved anatomy is an open problem rather than an immediate corollary."],"forward_implications":["Targeted virtual subgroups, such as older male cohorts, can be generated with within-group anatomical variability close to the real subgroup, which is what device and in silico trial studies need.","A trained CAN-FLOW can produce synthetic biventricular anatomies on demand without sharing individual image-derived meshes, easing privacy and data-access constraints.","The reported results imply that cVAE-style generators, despite closer nearest-neighbor fidelity, systematically underrepresent tail phenotypes and subgroup variability that matter for cohort-level predictions.","The conditional healthy-anatomy distribution could be used as a normative reference to score how far an individual heart deviates from expected shape for its sex, age, and BMI, pending validation in diseased cohorts."],"supporting_citations":[{"why":"supplies the automated segmentation and mesh-fitting pipeline that turns cardiac MR images into biventricular meshes.","marker":"[14]"},{"why":"provides the UK Biobank imaging cohort that defines the healthy population used for training and evaluation.","marker":"[21]"},{"why":"defines the LDDMM momenta preprocessing protocol and reports sex dimorphism in cardiac shape that the generated cohorts must reproduce.","marker":"[39]"},{"why":"introduces the LDDMM framework used to represent anatomies as diffeomorphic momenta with point-to-point correspondence.","marker":"[50]"},{"why":"provides the Glow-style flow architecture whose blocks CAN-FLOW adapts with conditional coupling and injector layers.","marker":"[54]"},{"why":"gives the variational autoencoder formulation that underlies the cVAE baseline and its ELBO objective.","marker":"[29]"},{"why":"gives the conditional VAE formulation used as the main baseline architecture.","marker":"[30]"},{"why":"introduces the coverage and minimum matching distance point-cloud metrics used to assess cohort spread and fidelity.","marker":"[43]"},{"why":"describes a prior conditional cardiac anatomy generator based on spatio-temporal generative modeling, providing cVAE baseline context.","marker":"[25]"},{"why":"describes a personalized time-resolved cardiac mesh generative model, another cVAE-style baseline context.","marker":"[26]"}],"fun_headline_variants":["Flow model beats VAEs at heart anatomy diversity","CAN-FLOW: two-step flow for realistic virtual heart cohorts","Normalizing flows capture more heart shape variation than VAEs","CAN-FLOW improves virtual heart cohort realism"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The evaluation assumes that comparing generated cohorts to the full 2,208-subject training cohort measures generalization; if cohort-level metrics are not restricted to the held-out test split, part of the reported distributional agreement could reflect the model matching anatomies it was trained on.","fun_headline_variants_meta":{"raw":{"variants":["Flow model beats VAEs at heart anatomy diversity","CAN-FLOW: two-step flow for realistic virtual heart cohorts","Normalizing flows capture more heart shape variation than VAEs","CAN-FLOW improves virtual heart cohort realism"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000674,"raw_usage":{"total_tokens":3089,"prompt_tokens":984,"completion_tokens":2105,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":600,"completion_tokens_details":{"reasoning_tokens":2040}},"tokens_in":600,"tokens_out":2105,"duration_ms":14626,"temperature":1.0,"reasoning_tokens":2040,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T17:06:07.519412+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute the KL divergences, Wasserstein distances, point-cloud coverage, and PCA variability ratios using only the 15% held-out test subjects as the real reference population. If CAN-FLOW's advantage over the cVAE baselines shrinks or reverses on out-of-training anatomies, the paper's central distributional-fidelity claim is refuted.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"supplies the automated segmentation and mesh-fitting pipeline that turns cardiac MR images into biventricular meshes."},{"cited_title":"Moscoloni, C","cited_arxiv_id":null,"evidence_quote":"defines the LDDMM momenta preprocessing protocol and reports sex dimorphism in cardiac shape that the generated cohorts must reproduce."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"gives the conditional VAE formulation used as the main baseline architecture."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"describes a prior conditional cardiac anatomy generator based on spatio-temporal generative modeling, providing cVAE baseline context."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"describes a personalized time-resolved cardiac mesh generative model, another cVAE-style baseline context."}],"review_version":1}