{"id":"c3096af3-bc75-4c5b-a6ac-3eed29009efe","arxiv_id":"2501.02146","paper_version":2,"verdict":"REJECT","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":5,"one_line_summary":"Plasma-CycleGAN conditions MRI-to-PET synthesis on plasma Aβ42/40 and reports improved similarity metrics, but the claimed consistent gains across all models are contradicted by the paper's own tables.","lead":"This paper adds a blood test result, the plasma Aβ42/40 ratio, as an extra input when generating brain PET images from MRI scans, and proposes a CycleGAN variant called Plasma-CycleGAN. The authors report quality improvements, but their main claim is contradicted by their own tables and the best model is chosen using the test set.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'consistent improvement' claim is contradicted by the paper's own Table 1 and is not backed by significance tests after test-set model selection.","rationale":"The reader's weakest assumption identifies lack of significance testing and test-set model selection, which is indeed load-bearing. I partially agree because the more immediate and concrete problem is the paper's own internal contradiction: the abstract's blanket claim of consistent enhancement is refuted by Table 1 for the image and add variants, even before considering statistics. The reader's rationale mentions this, but the weakest_assumption field focuses on selection bias. Both concerns are valid and independently justify rejection. The proposed concrete test would settle whether the concat-specific improvement is genuine: it would move from test-set peeking to validation-based selection and add paired significance testing. Given the paper's current state, the REJECT verdict stands unchanged.","tokens_in":6757,"tokens_out":4819,"duration_ms":48109,"concrete_test":"Re-run the experiment with a fixed random split; use only the 242-image validation set to choose the conditioning strategy (concat vs. none) for each baseline, then evaluate the selected models on the untouched 186-image test set. Report paired per-image SSIM differences (baseline vs. selected) with a Wilcoxon signed-rank test and a 95% bootstrap CI; for classification, report McNemar's test with subject-level clustering. If the concat gain is not significant at p<0.05 on the held-out test set, the 'consistent improvement' claim is unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that BBBM integration consistently improves generative quality is internally inconsistent: the abstract states it 'consistently enhances the generative quality across all models,' yet Table 1 shows Pix2pix+image (SSIM 0.683 vs 0.766), Pix2pix+add (0.465), and ShareGAN+image (0.556 vs 0.704) strongly degrade SSIM, MSE, and PSNR. Section 3.2 even admits 'the performance was unstable or decreased in Pix2pix and ShareGAN.' The narrower claim about concatenation surviving across all three baselines rests on differences of 0.014 (CycleGAN), 0.016 (Pix2pix), and 0.006 (ShareGAN) SSIM points, all far smaller than the reported standard deviations (~0.07–0.13). No paired significance tests are provided, and the concatenation strategy was selected after evaluating all 12 configurations on the test set (Table 1), so the differences may be artifacts of multiple testing and test-set peeking. The classification accuracy jump (CycleGAN 0.620 to Plasma-CycleGAN 0.815) is likewise reported without confidence intervals or a paired test, and the test set contains repeated images per subject, so the effective sample size is smaller than 186. These issues make the headline improvement unsubstantiated.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies whether incorporating the plasma Aβ42/40 ratio into three MRI-to-PET synthesis baselines (Pix2pix, CycleGAN, ShareGAN) improves generated PET quality. Three conditioning mechanisms are compared: adding the biomarker to the input image, adding it to a latent feature map, and concatenating it as an extra channel in latent space. The paper reports that the concatenation strategy consistently improves all three baselines, names the CycleGAN variant Plasma-CycleGAN, and further evaluates SUVR correlation and amyloid-positivity classification on the synthesized PET images. The authors claim that this is the first integration of blood-based biomarkers into MRI-to-PET cross-modality translation.","tokens_in":7032,"tokens_out":6352,"duration_ms":58351,"significance":"If the claims were substantiated, the paper would introduce a clinically meaningful, low-cost conditioning variable into cross-modality image translation, with potential diagnostic utility. The use of a moderately large ADNI dataset, multiple baselines, and a subject-level train/validation/test split is a reasonable experimental setup. The clearest strength is that the idea is concrete and the reported numbers are falsifiable. However, the paper's central claim as written is contradicted by its own Table 1, and the remaining positive results are presented without significance testing or proper model selection, so the practical significance of the proposed method is not currently established.","major_comments":[{"comment":"The claim that \"BBBMs integration consistently enhances the generative quality across all models\" is contradicted by the paper's own results. In Table 1, Pix2pix+image has SSIM 0.683 vs. baseline 0.766, Pix2pix+add has SSIM 0.465 vs. 0.766, and ShareGAN+image has SSIM 0.556 vs. 0.704; these are substantial degradations, not enhancements. The Table 1 caption stating \"All models achieved improved performance after incorporating BBBMs in all metrics\" is factually incorrect for these rows. The text in Sec. 3.2 later admits that \"the performance was unstable or decreased in Pix2pix and ShareGAN.\" The abstract, introduction, and conclusion should be corrected to state the narrower, accurate finding that latent-space concatenation produced small directional improvements.","section":"Abstract and Sec. 3.2, Table 1"},{"comment":"The surviving claim that latent-space concatenation improves all three baselines is not statistically supported. The SSIM differences are +0.014 (CycleGAN), +0.016 (Pix2pix), and +0.006 (ShareGAN), while the reported standard deviations are on the order of 0.07–0.13. The PSNR differences are +0.42, +0.34, and +0.05 dB, and the MSE differences are also small relative to the standard deviations. No paired significance tests, confidence intervals, or effect-size statistics are provided, even though paired tests on the 186 test images should be straightforward. Without such tests, the observed differences are within the noise level and do not support the claim of consistent improvement.","section":"Sec. 3.2, Table 1"},{"comment":"The model selection procedure is problematic. Table 1 reports test-set performance for all 12 configurations, and Sec. 3.2 states \"we only use concatenation in later experiments\" immediately after presenting those test-set results. This constitutes selecting the best variant based on the same test set used for the final evaluation, which inflates the reported performance and makes the subsequent Plasma-CycleGAN comparisons (Tables 2 and 3) difficult to interpret. No validation-based selection, hold-out procedure, or correction for multiple comparisons is described. A proper approach would be to select the integration method on the 242-image validation set and to report test-set results only for the selected configuration.","section":"Sec. 3.1 and Sec. 3.2"},{"comment":"The classification accuracy improvement (CycleGAN 0.620 to Plasma-CycleGAN 0.815) is reported without confidence intervals, significance tests, or a description of the number of unique subjects in the test set. Because the data consist of 1338 images from 456 individuals and the split is by subject, the 186 test images are not independent; multiple images can come from the same person. The effective sample size for classification is therefore smaller than 186, and treating images as independent observations inflates the apparent statistical strength. A per-subject bootstrap, a cluster-robust test, or reporting subject-level aggregated predictions is needed to support the claim that Plasma-CycleGAN achieves the best accuracy.","section":"Sec. 3.3, Table 3"},{"comment":"The SUVR correlation results are likewise reported without accounting for the clustered structure of the data. Pearson correlation coefficients and p-values computed on 186 images from a smaller number of subjects do not satisfy the independence assumption. The claim that \"incorporating BBBM information enhanced the correlation in all models\" should be reassessed with a cluster-robust correlation or a per-subject analysis. The differences between PCC values (e.g., 0.777 to 0.807 for CycleGAN) are also presented without confidence intervals, so it is unclear whether they are meaningful.","section":"Sec. 3.3, Table 2"}],"minor_comments":[{"comment":"The caption \"All models achieved improved performance after incorporating BBBMs in all metrics\" is inaccurate and should be revised to reflect the actual rows, several of which show degraded SSIM, PSNR, or MSE.","section":"Table 1"},{"comment":"There is a typographical error in the Pix2pix+add cell: \"259.09.01±271.76\" should likely be \"259.09±271.76\" or a corrected value.","section":"Table 1"},{"comment":"The loss function for CycleGAN is written with inconsistent notation: after defining L(G1,G2,D1,D2), the identity loss term is written as \"λidtLidt\" without defining Lidt; please clarify whether this is the same as Lide used for ShareGAN.","section":"Sec. 2.3"},{"comment":"The statement that \"CycleGAN showed the highest robustness against disturbance\" is not supported by any experiment or quantitative criterion; if robustness is intended as a claim, it should be formalized and measured.","section":"Sec. 3.2"},{"comment":"The sentence \"The difference of validation and testing sets was to avoid data leakage\" is confusing; splitting by subject avoids leakage, whereas validation and test sets serve the standard model-selection and final-evaluation purposes. Please clarify the intended roles of the validation and test subsets.","section":"Sec. 3.1"}],"recommendation":"major_revision","confidential_remarks":"The manuscript has a good overall motivation and a reasonable experimental skeleton, but the central claim is overstated and at least one load-bearing result is internally inconsistent. The authors should be asked to correct the abstract and conclusions to reflect the actual (narrower) finding, and to add statistical significance tests and a proper model-selection protocol. If the authors cannot provide such evidence for the concat variant after revision, the paper may not meet the bar for publication. The manuscript could be suitable for a venue that values the empirical systematic comparison, provided the text is made honest with respect to what the data actually show."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The genuinely new thing here is conditioning MRI-to-PET synthesis on plasma Aβ42/40 and comparing three injection mechanisms (add to input, add to latent, concat to latent) across three GAN baselines. That is a legitimate, modest extension, and the concat method does improve SSIM for all three baselines. The classification jump for CycleGAN with concat (0.620 to 0.815) is also worth taking seriously, though it comes with caveats. Credit where due: the authors test a clinically motivated covariate in a straightforward way and report their numbers honestly in the table, even if the summary statements misrepresent them.\n\nThe load-bearing flaw is the headline claim. The abstract and the Table 1 caption say BBBM integration \"consistently enhances\" generative quality across all models, but Table 1 shows Pix2pix+image and ShareGAN+image are clearly worse than their baselines on SSIM, PSNR, and MSE, and Pix2pix+add drops SSIM from 0.766 to 0.465. The body text does admit instability for Pix2pix and ShareGAN, so the abstract overstates the case. The claim only survives for concat, and the improvements there are small (SSIM +0.014, +0.016, +0.006) relative to the reported standard deviations (0.07–0.13). No significance tests are provided. More concerning, the winning concat variant was selected after evaluating all 12 configurations on the test set, so the improvement may be a multiple-comparisons artifact. The classification accuracy difference also lacks confidence intervals or a paired test, and the test set likely contains repeated images per subject, making the effective sample size smaller than 186. No code or data are released, which limits reproducibility.\n\nThere is also a clear typo in Table 1 (\"259.09.01\") that suggests a rushed edit. That is minor by itself, but combined with the internal contradiction it undercuts confidence in the manuscript's polish.\n\nThe core idea is not wrong; the paper is just oversold. A revised version that corrects the claims, selects the model on validation, reports significance tests, and ideally releases code could be a reasonable contribution. As is, I would not cite it or trust the headline result. But it deserves a serious referee rather than a desk reject, because the topic is relevant, the approach is understandable, and the flaws are addressable rather than fatal.","headline":"A clinically motivated but overclaimed GAN conditioning study; the concat variant may help, but the headline claim is contradicted by the paper's own Table 1 and lacks statistical support.","tokens_in":7567,"tokens_out":3616,"would_cite":false,"duration_ms":33819,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Conditioning MRI-to-PET synthesis on the plasma Aβ42/40 biomarker consistently improves generated PET image quality and amyloid-positivity classification; the best configuration, CycleGAN with latent concatenation, is named Plasma-CycleGAN.","keywords":["MRI to PET synthesis","CycleGAN","blood-based biomarkers","plasma Aβ42/40","Alzheimer's disease","cross-modality translation","amyloid PET","conditional image generation"],"falsifier":"Re-run the comparison with a validation-based selection procedure: pick the conditioning method on a held-out validation set, then compare Plasma-CycleGAN with plain CycleGAN on a test set used only once, and repeat the split to estimate uncertainty; if the SSIM or classification gap disappears, the paper's central claim is not supported.","tokens_in":6566,"feed_emoji":"🧠","tokens_out":6937,"duration_ms":62770,"temperature":0.7,"pith_summary":"This paper tries to establish that a blood-test value—the plasma Aβ42/40 ratio—can serve as an explicit condition in MRI-to-PET synthesis for Alzheimer's disease, making the synthesized PET images better. Across three generative baselines, the authors report that injecting the biomarker into the latent space by concatenation consistently improves image-similarity metrics, and that the best overall model, Plasma-CycleGAN, also classifies amyloid positivity more accurately than the same model without the biomarker. If the result holds, inexpensive blood biomarkers could guide image translation and make synthetic PET more clinically informative without an actual PET scan.","feed_headline":"Blood biomarker lifts AI-generated PET image quality","feed_subtitle":"Conditioning MRI-to-PET translation on plasma Aβ42/40 improves fidelity and amyloid classification in the paper's tests.","key_machinery":"The load-bearing mechanism is latent-space concatenation of the biomarker: a scalar normalized plasma Aβ42/40 value is expanded into a single-channel $16 \\times 16 \\times 16$ tensor and concatenated with the generator's $128 \\times 16 \\times 16 \\times 16$ bottleneck feature map, yielding a 129-channel map that a $1 \\times 1$ convolution reduces back to 128 channels. This injects the biological information precisely where the encoder's structural features are about to be decoded into PET, and the paper reports that this placement—unlike adding the value to the input image or to the feature map—improves performance consistently across Pix2pix, CycleGAN, and ShareGAN.","core_discovery":"The central claim of the paper is that incorporating blood-based biomarker information, specifically the normalized plasma Aβ42/40 ratio, into deep generative models improves MRI-to-PET translation. The paper tests three conditioning placements for each of three baselines (Pix2pix, CycleGAN, ShareGAN) and finds that expanding the scalar biomarker to the size of the bottleneck feature map and concatenating it as an extra channel improves SSIM, PSNR, and MSE for every baseline, while the other two placements are unstable or harmful. For CycleGAN the reported gain is from SSIM 0.808 to 0.822, PSNR 24.65 to 25.07, and MSE 250.33 to 227.00, and amyloid-positivity classification accuracy jumps from 0.620 to 0.815. The paper therefore names CycleGAN plus latent concatenation Plasma-CycleGAN and presents it as the first cross-modality MRI-to-PET translation model conditioned on blood-based biomarkers.","pith_inferences":["A natural next experiment, which the paper does not run, is to test whether latent concatenation of plasma Aβ42/40 also improves diffusion-based MRI-to-PET synthesis; the consistency across three GAN baselines is suggestive but untested for diffusion models.","If the central claim holds, a blood draw and a routine MRI could in principle serve as a PET surrogate for amyloid assessment, making amyloid screening more accessible; this clinical workflow is not evaluated in the paper.","The injection mechanism is generic to scalar clinical variables, so the same design could be tried with p-tau217 or other biomarkers; the paper names p-tau217 as future work but does not test it."],"forward_implications":["If the reported numbers are representative, synthetic PETs produced by Plasma-CycleGAN would carry enough amyloid signal to separate amyloid-positive from amyloid-negative brains with 0.815 accuracy, compared with 0.620 for the unconditioned CycleGAN.","The fact that latent concatenation helped all three baselines suggests the gain comes from where the biomarker enters the network, not from a particular generator architecture.","Improved SSIM, PSNR, and SUVR correlation together imply the generated images are closer to real PET both pixel-wise and in the uptake-value distribution used clinically.","The paper's stated next steps are to apply the same conditioning idea to diffusion models and to incorporate p-tau217, a biomarker it notes has shown accuracy comparable to cerebrospinal fluid measures."],"supporting_citations":[{"why":"Supplies the CycleGAN architecture and cycle-consistency loss that Plasma-CycleGAN extends with biomarker conditioning.","marker":"[7]"},{"why":"Supplies the Pix2pix conditional-GAN baseline and its objective, which the study compares against CycleGAN.","marker":"[15]"},{"why":"Supplies the ShareGAN baseline and the modified discriminator architecture used in the CycleGAN implementation.","marker":"[16]"},{"why":"Provides the MCSUVR threshold used to define amyloid positivity in the classification evaluation.","marker":"[14]"},{"why":"Establishes plasma Aβ42/40 as a relevant blood-based biomarker for brain amyloid, motivating the conditioning variable.","marker":"[11]"},{"why":"Supports the claim that blood-based biomarkers, including Aβ42/40, have strong potential for detecting brain amyloid.","marker":"[12]"},{"why":"Provides clinical context for blood-based biomarkers as a minimally invasive route to amyloid assessment.","marker":"[13]"}],"fun_headline_variants":["Plasma biomarker boosts MRI-to-PET image synthesis","CycleGAN + blood biomarker sharpens synthetic PET","Blood protein improves AI PET fidelity and amyloid read","Biomarker-conditioned CycleGAN enhances MRI-to-PET translation"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole result rests on the assumption that the reported improvement is a real effect of adding the blood-test value and not just a lucky pick among the twelve models the authors compared.","fun_headline_variants_meta":{"raw":{"variants":["Plasma biomarker boosts MRI-to-PET image synthesis","CycleGAN + blood biomarker sharpens synthetic PET","Blood protein improves AI PET fidelity and amyloid read","Biomarker-conditioned CycleGAN enhances MRI-to-PET translation"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00075,"raw_usage":{"total_tokens":3323,"prompt_tokens":911,"completion_tokens":2412,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":527,"completion_tokens_details":{"reasoning_tokens":2348}},"tokens_in":527,"tokens_out":2412,"duration_ms":19545,"temperature":1.0,"reasoning_tokens":2348,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T22:14:02.419300+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the comparison with a validation-based selection procedure: pick the conditioning method on a held-out validation set, then compare Plasma-CycleGAN with plain CycleGAN on a test set used only once, and repeat the split to estimate uncertainty; if the SSIM or classification gap disappears, the paper's central claim is not supported.","supporting_citations":[{"cited_title":"A/T/N: An unbiased descriptive classification scheme for Alzheimer disease biomarkers,","cited_arxiv_id":null,"evidence_quote":"Supplies the CycleGAN architecture and cycle-consistency loss that Plasma-CycleGAN extends with biomarker conditioning."},{"cited_title":"FREA-Unet: Frequency-aware U-net for Modality Transfer","cited_arxiv_id":"2012.15397","evidence_quote":"Supplies the Pix2pix conditional-GAN baseline and its objective, which the study compares against CycleGAN."},{"cited_title":"PASTA: Pathology-Aware MRI to PET Cross-Modal Translation with Diffusion Models","cited_arxiv_id":"2405.16942","evidence_quote":"Provides the MCSUVR threshold used to define amyloid positivity in the classification evaluation."},{"cited_title":"Bidirectional map- ping generative adversarial networks for brain MR to PET synthesis,","cited_arxiv_id":null,"evidence_quote":"Establishes plasma Aβ42/40 as a relevant blood-based biomarker for brain amyloid, motivating the conditioning variable."},{"cited_title":"MRI to PET Cross-Modality Translation using Globally and Locally Aware GAN (GLA-GAN) for Multi-Modal Diagnosis of Alzheimer's Disease","cited_arxiv_id":"2108.02160","evidence_quote":"Supports the claim that blood-based biomarkers, including Aβ42/40, have strong potential for detecting brain amyloid."},{"cited_title":"Unpaired image-to-image translation using cycle-consistent adversarial networks,","cited_arxiv_id":null,"evidence_quote":"Provides clinical context for blood-based biomarkers as a minimally invasive route to amyloid assessment."}],"review_version":1}