{"id":"d9826933-c069-4817-a509-66dca84b9b2e","arxiv_id":"1908.08074","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"DUAL-GLOW synthesizes PET brain images from MRI using two flow networks and a latent relation network, with optional age conditioning, and shows strong quantitative results on ADNI.","lead":"This paper describes DUAL-GLOW, a generative model that creates synthetic PET brain scans from MRI scans by learning the relationship between the two image types in a compressed space. It also shows how conditioning on age makes the model display expected declines in brain metabolism, with potential uses in Alzheimer's research.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Table 1 does not statistically substantiate the claim of quantitative superiority: DUAL-GLOW loses on CorCoef and its PSNR/SSIM gains overlap with baseline variability, so 'quantitatively better than recent works' is unsupported as stated.","rationale":"DUAL-GLOW's core derivation is a valid conditional normalizing-flow construction; Eq. (10) follows from the change-of-variables rule, and Eq. (16) is a reasonable regularized objective. The availability of source code is helpful, although the reader could not execute it. I did not make the Gaussian-latent conditional the primary attack: since fp is an invertible flow, p(xp|xm) is a warped Gaussian in latent space and can represent a fairly wide family; without a goodness-of-fit test or direct comparison to a more flexible conditional flow, the Gaussianity concern is not decisive. The load-bearing weakness is that the empirical claim of superiority is not statistically established and one headline metric contradicts it. This matches the reader's conditional verdict, so no change in verdict is needed.","tokens_in":14769,"tokens_out":12287,"duration_ms":136344,"concrete_test":"Reconstruct the exact 10-fold or 10-split protocol from Section 4.1 and run paired permutation or bootstrap tests on per-fold or per-subject metrics from Table 1, comparing DUAL-GLOW against C-VAE and pix2pix on MAE, CorCoef, PSNR, and SSIM. Report 95% confidence intervals for each metric difference. If the CI for PSNR or SSIM includes zero, or if CorCoef remains significantly worse, weaken the 'quantitatively better' claim to parity.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The mathematical derivation in Eqs. (8)-(12) is internally sound: with invertible flows fm and fp, the conditional log-likelihood factorizes into a latent conditional, a PET Jacobian, and a regularized MR marginal. My concern is the empirical payoff claimed in the abstract and Section 4.2. Table 1 reports DUAL-GLOW CorCoef 0.975 versus C-VAE 0.980, so on one of the four headline metrics DUAL-GLOW is numerically worse. For PSNR, the reported mean±std values are 29.56±2.66 (DUAL-GLOW) and 28.69±2.06 (C-VAE); the one-standard-deviation intervals overlap substantially, and no significance tests or confidence intervals are reported for any metric. Section 4.1 describes 'randomly select 726 subjects as training data and the remaining 80 as testing within a 10-fold evaluation scheme,' which is not a standard 10-fold cross-validation and leaves the effective number of independent test sets unclear. The age-conditioning evidence in Figure 7 carries the same statistical weakness; the authors themselves note that wide variance bands mean 'a larger sample size may be necessary to derive statistically sound conclusions.' Despite this, the abstract states 'quantitatively better than recent works' and that the model can 'capture brain FDG-PET changes, as a function of age' without the same caveat. If the observed PSNR/SSIM advantages are within fold-to-fold noise and CorCoef stays below the C-VAE baseline, the central claim reduces to parity with existing methods, not superiority.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces DUAL-GLOW, a normalizing-flow model for conditional MRI-to-PET synthesis. Two invertible networks map MRI and PET into latent spaces, and a relation network models the conditional density of the PET latent given the MRI latent as a Gaussian whose mean and variance are network outputs. The conditional log-likelihood is derived via the change-of-variables formula in Eqs. (8)-(12). The framework is extended to side information such as age by adding discriminators and gradient reversal layers intended to confine the effect of the covariate to the top-most latent level. Experiments on ADNI compare DUAL-GLOW with cGAN, UcGAN, C-VAE, and pix2pix using MAE, CorCoef, PSNR, and SSIM, and include an SVM classification of generated images and age-conditioned PET generation.","tokens_in":15096,"tokens_out":6078,"duration_ms":55774,"significance":"If the claims hold, the paper provides a useful likelihood-based alternative to GANs for medical modality transfer, with the practical advantage of exact latent inference and full 3D-volume processing. The mathematical derivation in Section 3 is clean and self-contained: Eq. (10) correctly follows from the block-diagonal Jacobian in Eq. (11), and Eq. (12) is a valid conditional-likelihood objective with a regularizer on the MRI marginal. The release of code is a concrete strength. However, the empirical support for the headline superiority claim is incomplete: no significance tests are reported, the cross-validation protocol is ambiguous, one of the four headline metrics (CorCoef) is below the C-VAE baseline, and the age-conditioning results are qualitative with wide variance bands. The significance of the paper is therefore moderate pending strengthened statistical evidence and a more measured abstract.","major_comments":[{"comment":"The abstract's claim that DUAL-GLOW is 'quantitatively better than recent works' is not supported by the reported statistics. DUAL-GLOW has CorCoef 0.975 versus C-VAE 0.980, so it is numerically worse on one of the four headline metrics. For PSNR, the means are 29.56±2.66 versus 28.69±2.06, and for SSIM 0.898±0.06 versus 0.817±0.06; the one-standard-deviation intervals for PSNR overlap substantially, and no confidence intervals, paired tests, or per-fold results are provided. Please add a properly paired significance analysis across the folds or subjects with multiple-comparison control and report effect sizes, or revise the abstract to 'comparable or better' with appropriate caveats.","section":"Section 4.2, Table 1"},{"comment":"The sentence 'randomly select 726 subjects as the training data and the remaining 80 as testing within a 10-fold evaluation scheme' is ambiguous. A standard 10-fold cross-validation on 806 subjects would use test sets of roughly 80 subjects, but a single random 80-subject split is not 10-fold. Please specify exactly how many folds or splits were used, whether the same folds were used for all compared methods, and how test subjects were selected for the age-conditioning experiments; this determines the effective number of independent test evaluations and the validity of any significance statements.","section":"Section 4.1"},{"comment":"The age-conditioning contribution rests on the assumption that side information affects only the top-most latent level and that gradient reversal removes age from the lower levels. No quantitative evidence is given that the lower-level latents are indeed age-invariant, for example via classifier accuracy on those latents before and after the gradient reversal layers. The only quantitative support is Figure 7, whose 95% bands are described by the authors as too wide for statistically sound conclusions. The abstract's statement that the model can 'capture brain FDG-PET changes as a function of age' should therefore be backed by a statistical test of the age trend with confidence intervals and a validation of the disentanglement assumption, or softened to a qualitative claim.","section":"Section 3, 'How to condition based on side information'; Figure 7"},{"comment":"The objective in Eq. (12) is exactly the conditional log-likelihood only under the Gaussian family pθ(zp|zm)=N(zp;µθ(zm),σθ(zm)). The paper provides no diagnostic for this assumption. Please include a residual analysis on held-out subjects, for example standardized residuals (zp−µθ(zm))/σθ(zm) compared with a standard normal, or a comparison against a more flexible conditional density, or explicitly discuss the limitation. This is a correctness-risk check on the central likelihood interpretation rather than a request for a different model.","section":"Section 3, Eq. (6)"}],"minor_comments":[{"comment":"The dataset size is given as 826 subjects in the abstract but 806 clean MRI/PET pairs in Section 4.1; please reconcile these numbers.","section":"Abstract and Section 4.1"},{"comment":"The optimizer is called 'AdamMax' but should be 'AdaMax' to match the terminology in reference [20].","section":"Section 4.1"},{"comment":"The caption contains a typo: 'Spliting' should be 'Splitting'.","section":"Figure 3 caption"},{"comment":"The metric name 'Cor Coef' should be written consistently as 'CorCoef' as in Table 1.","section":"Section 4.2"},{"comment":"The x-axis is labeled only '50 100'; please label the axis 'Age' with explicit tick values and add units.","section":"Figure 7"},{"comment":"The notation I1:d1 is not defined; it should denote the identity matrix acting on the first d1 components.","section":"Equation (15)"}],"recommendation":"major_revision","confidential_remarks":"The paper is within the journal's scope and the mathematical core is sound. My main reservation is the gap between the strong abstract claims and the statistical support; this is fixable with additional analysis or revised wording. The Gaussian-conditional assumption is a modeling choice common in the flow literature, so I would not reject solely on that ground if residual diagnostics are supplied."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The math is sound and the framework is worth knowing about, but the quantitative superiority claim over baselines is not actually backed by the numbers the paper itself reports. Two invertible flows with a learned Gaussian relation in latent space is a legitimate extension of GLOW for paired modality transfer, and the derivation in Eqs. (8)-(12) is correct: the conditional log-likelihood factorizes into the latent conditional plus the PET Jacobian, and the block-matrix determinant gives the clean form. The side-information handling via GRL-based disentanglement is also a decent idea, even if the assumption that only the top level carries age is stated rather than tested.\n\nWhere the paper gets soft is the experiments. Table 1 shows CorCoef of 0.975 for DUAL-GLOW against 0.980 for C-VAE, so they lose on one of four headline metrics. On PSNR the lead is 29.56±2.66 versus 28.69±2.06; the one-standard-deviation intervals overlap, and there are no significance tests or confidence intervals anywhere. SSIM is a clearer win (0.898 vs 0.817) but the paper still calls the overall result \"quantitatively better\" without statistical support. The 10-fold evaluation is also described ambiguously: \"randomly select 726 subjects as training and the remaining 80 as testing within a 10-fold evaluation scheme\" is not standard cross-validation, and the effective number of independent test sets is unclear.\n\nThe age-conditioning part is weaker than it looks. The downward ROI trend is a direct product of conditioning on age, not an independent prediction, and the paper itself admits wide variance bands and the need for a larger sample. Yet the abstract states the model can \"capture brain FDG-PET changes as a function of age\" without that caveat. That mismatch should be fixed.\n\nCredit where due: the downstream classification test on generated PET is a nice idea, the figures look plausible, and the code is available on GitHub, though I did not execute it. The central claim—that a flow-based conditional model can produce meaningful PET from MRI in a small-sample regime—is credible and worth publishing. But the paper needs a statistical revision: significance tests or CIs, a clearer description of the split, and a more measured abstract. This would be a useful contribution for researchers working on flow-based medical image synthesis, and it deserves serious peer review rather than a desk reject.","headline":"The architectural idea and derivation are solid, but the empirical headline overreaches: the numbers in Table 1 don't support 'quantitatively better' without significance tests.","tokens_in":15645,"tokens_out":1945,"would_cite":true,"duration_ms":19779,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"DUAL-GLOW is a flow-based generative model that learns the conditional distribution of PET given MRI and reports synthetic 3D volumes that outperform GAN-based baselines on the ADNI dataset.","keywords":["flow-based generative model","modality transfer","MRI-to-PET synthesis","normalizing flows","conditional likelihood","Alzheimer's disease","FDG-PET","side information conditioning"],"falsifier":"Train an age predictor on the lower-level MRI latent codes of the age-conditioned model; if its accuracy is well above chance, the top-level-only assumption is violated, and similarly, replacing the Gaussian conditional with a more flexible density and showing materially better held-out conditional log-likelihood would falsify the Gaussianity assumption.","tokens_in":14554,"feed_emoji":"🧠","tokens_out":6875,"duration_ms":66001,"temperature":0.7,"pith_summary":"The paper tries to establish that flow-based generative models can perform MRI-to-PET modality transfer in a small-sample medical setting by learning the conditional distribution of PET given MRI in latent space. It argues that with two invertible networks and a relation network, maximizing conditional likelihood reduces to a tractable objective and yields sharper, more faithful synthetic PET than adversarial and autoencoder baselines. It also claims that side information such as age can be disentangled into the top latent level, so the model can generate PET images that reflect age-related hypometabolism. If these claims hold, PET synthesis becomes a likelihood-based, invertible procedure with exact latent inference and useful downstream diagnostic signal.","feed_headline":"Flow-based DUAL-GLOW turns brain MRI into PET better than GANs","feed_subtitle":"Two invertible networks plus a latent relation net produce sharper 3D PET and track age-related hypometabolism.","key_machinery":"The load-bearing object is a paired-flow likelihood: two invertible affine-coupling flows with multi-scale splitting, and a relation network that parameterizes the Gaussian conditional density p(z_p|z_m) = N(z_p; mu_theta(z_m), sigma_theta(z_m)). The identity log p(x_p|x_m) = log p(z_p|z_m) + log|det(dz_p/dx_p)| lets the model train with exact log-likelihood; affine coupling makes the Jacobian log-determinant a simple sum over scale terms, and multi-scale splitting reduces computation. For side information, gradient-reversal-layer-equipped discriminators strip age or attribute signal from lower latent levels while a top-level discriminator preserves it, enforcing the assumption that only the highest level of the representation is affected by the conditioning variable.","core_discovery":"DUAL-GLOW's central claim is that the conditional distribution of PET given MRI can be learned with two normalizing flows, one mapping PET to a latent code and one mapping MRI to its own latent code, plus a relation network that predicts a Gaussian conditional density between the two latent spaces. Under the change-of-variables rule, the conditional log-likelihood of PET given MRI equals the log of this latent conditional density plus the log-determinant of the PET flow's Jacobian, and adding a regularizer on the MRI marginal gives the training objective in Eq. (12). The paper reports that on 806 ADNI MRI/PET pairs, DUAL-GLOW produces full 3D PET volumes with higher SSIM and PSNR and lower MAE than cGAN, UcGAN, C-VAE and pix2pix, and that an SVM trained on synthetic PET achieves AD/CN classification accuracy comparable to ground truth. The age-conditioned extension shows decreasing regional intensity with age, matching the expected aging-related hypometabolism.","pith_inferences":["Because the two flows are invertible, the framework could in principle run in reverse to estimate plausible MRI volumes from PET, a direction the paper does not test.","The Gaussian conditional assumption is not directly validated; an immediate extension would be replacing p(z_p|z_m) with a conditional normalizing flow on the latent pair and measuring whether held-out conditional log-likelihood or synthesis quality improves.","The GRL-disentanglement design could be applied to other side variables such as disease status, sex, or genotype, though the paper's wide variance bands suggest larger sample sizes are needed for reliable group comparisons.","The relation network's Gaussian parameters provide a natural source of uncertainty: repeated sampling from p(z_p|z_m) could yield voxel-level confidence intervals for metabolism, which the paper does not report."],"forward_implications":["If DUAL-GLOW is correct, MRI-to-PET synthesis can be trained by maximizing a tractable conditional likelihood rather than by adversarial or reconstruction losses, which avoids the blur and mode collapse associated with those approaches.","Because both transformations are invertible, the same trained pair of flows provides exact latent-variable inference and lets a practitioner sample diverse PET volumes for a single MRI rather than one deterministic output.","The age-conditioned extension implies that, for a fixed MRI, PET-like hypometabolism changes with the side variable can be simulated, giving a tool for studying progression of neurodegeneration across the age range.","Downstream diagnostic classifiers trained on synthetic PET retain most of the discriminative signal of real PET, with the paper reporting 91% versus 94% accuracy for AD/CN classification, so generated volumes could augment small cohorts in statistical analyses."],"supporting_citations":[{"why":"Supplies GLOW, the base flow-based generative model whose multi-scale architecture and affine coupling layers DUAL-GLOW extends to two parallel invertible networks.","marker":"[21]"},{"why":"Provides the affine coupling layer and multi-scale splitting technique used to build the invertible functions f_m and f_p and to keep Jacobian log-determinants tractable.","marker":"[8]"},{"why":"Establishes the normalizing-flow change-of-variable framework that underlies the conditional log-likelihood derivation.","marker":"[35]"},{"why":"Gives the gradient reversal layer used to strip side information from lower latent levels in the age-conditioned model.","marker":"[10]"},{"why":"Justifies the block-determinant factorization used to reduce the joint Jacobian to det(dz_p/dx_p) in the derivation of Eq. (10).","marker":"[38]"},{"why":"Is the pix2pix conditional adversarial baseline that DUAL-GLOW compares against on ADNI and on natural image translation.","marker":"[16]"},{"why":"Provides the context-aware GAN baseline for medical image synthesis used in the quantitative comparison.","marker":"[32]"},{"why":"Defines the U-Net architecture used in the UcGAN baseline and in autoencoder-based synthesis approaches.","marker":"[36]"},{"why":"Supplies the atlas-based segmentation that yields the 116 ROI features used for the downstream AD/CN classification of synthetic PET.","marker":"[45]"}],"fun_headline_variants":["Flow-based MRI-to-PET generator beats GANs on small data","DUAL-GLOW: invertible nets turn MRI into PET, rivaling GANs","Conditional flow model synthesizes PET from MRI, outshines GANs","Age-conditioned PET synthesis from MRI with flow-based DUAL-GLOW","MRI to PET: flow-based model outperforms GANs in small-sample regime"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole derivation assumes that, after the two flows, the conditional density of the PET latent given the MRI latent is exactly Gaussian with mean and variance produced by neural networks, and the age-conditioning version further assumes side information lives only in the top latent level; the paper gives no direct evidence for either assumption.","fun_headline_variants_meta":{"raw":{"variants":["Flow-based MRI-to-PET generator beats GANs on small data","DUAL-GLOW: invertible nets turn MRI into PET, rivaling GANs","Conditional flow model synthesizes PET from MRI, outshines GANs","Age-conditioned PET synthesis from MRI with flow-based DUAL-GLOW","MRI to PET: flow-based model outperforms GANs in small-sample regime"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000213,"raw_usage":{"total_tokens":1445,"prompt_tokens":992,"completion_tokens":453,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":608,"completion_tokens_details":{"reasoning_tokens":347}},"tokens_in":608,"tokens_out":453,"duration_ms":3969,"temperature":1.0,"reasoning_tokens":347,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T11:50:37.387489+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train an age predictor on the lower-level MRI latent codes of the age-conditioned model; if its accuracy is well above chance, the top-level-only assumption is violated, and similarly, replacing the Gaussian conditional with a more flexible density and showing materially better held-out conditional log-likelihood would falsify the Gaussianity assumption.","supporting_citations":[{"cited_title":"Glow: Generative ﬂow with invertible 1x1 convolutions","cited_arxiv_id":null,"evidence_quote":"Supplies GLOW, the base flow-based generative model whose multi-scale architecture and affine coupling layers DUAL-GLOW extends to two parallel invertible networks."},{"cited_title":"Determinants of block matrices","cited_arxiv_id":null,"evidence_quote":"Justifies the block-determinant factorization used to reduce the joint Jacobian to det(dz_p/dx_p) in the derivation of Eq. (10)."},{"cited_title":"Image-to-image translation with conditional adversar- ial networks","cited_arxiv_id":null,"evidence_quote":"Is the pix2pix conditional adversarial baseline that DUAL-GLOW compares against on ADNI and on natural image translation."},{"cited_title":"Medical image syn- thesis with context-aware generative adversarial networks","cited_arxiv_id":null,"evidence_quote":"Provides the context-aware GAN baseline for medical image synthesis used in the quantitative comparison."},{"cited_title":"Opti- mum template selection for atlas-based segmentation","cited_arxiv_id":null,"evidence_quote":"Supplies the atlas-based segmentation that yields the 116 ROI features used for the downstream AD/CN classification of synthetic PET."}],"review_version":1}