{"id":"27f985af-c312-4676-a64c-5f66c95126a7","arxiv_id":"2506.11297","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"In 52 epilepsy patients, score-based diffusion models synthesized full-dose FDG brain PET from MRI with lower voxel error than a Transformer-Unet, and 1% dose PET inputs made all models comparable.","lead":"Researchers compared two score-based diffusion models and a Transformer-Unet for turning brain MRI (with or without 1% dose PET) into full-dose FDG brain PET in epilepsy patients. The best diffusion model matched or beat the transformer when only MRI was available, and adding ultra-low-dose PET made all three models perform similarly.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Clinical 'accurate synthesis' claim is not established: CI/CMAE only test sign and area-weighted errors, and can miss small focal hypometabolism; with 10 subjects and no reader study, similarity to acquired PET does not prove diagnostic utility.","rationale":"The most defensible reading of the paper is a feasibility comparison of three generative models on a small single-center epilepsy PET/MRI dataset, with honest acknowledgement that reader studies remain future work. The engineering contribution is plausible and should not be rejected. However, the abstract and conclusion make a clinical accuracy claim that outruns the evidence: CI/CMAE and SUVR/ICC are similarity metrics, and CI/CMAE in particular can be satisfied by a synthetic PET that misses the small focal abnormality that matters for epilepsy surgery. That makes the central claim load-bearing on a validation that is absent. The reader already identified the lack of reader-study validation; this stress-test extends it by showing the specific mechanism (sign-only CI, area-weighted CMAE) by which the proposed metrics can be misleading. Since the paper itself qualifies these metrics as needing future correlation with reader studies, and since the concern is about overclaiming rather than an internal error, the existing CONDITIONAL verdict remains appropriate.","tokens_in":14728,"tokens_out":7758,"duration_ms":92961,"concrete_test":"Run a blinded localization study on the 10 test patients (and ideally an independent MRI-negative epilepsy cohort): have 2-3 nuclear medicine physicians lateralize and localize the epileptogenic zone from (a) acquired full-dose PET, (b) SGM-KD zero-dose synthetic PET, and (c) the best ultralow-dose synthetic PET, and compare with the clinical gold standard (SEEG, surgical outcome, or consensus clinical report). Report lobe-level sensitivity/specificity and laterality accuracy for each; if zero-dose synthetic PET's localization accuracy is substantially below acquired PET, the claim that MRI-only synthesis 'synthesizes full-dose FDG-PET accurately' fails regardless of CI/CMAE.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing step is the inference from image-similarity metrics to the abstract/conclusion claim that all three models can 'synthesize full-dose FDG-PET accurately'. The metrics computed against the acquired PET do not establish clinical accuracy. In Eq. (15), the Congruence Index counts an ROI as matching whenever the sign of the synthetic and acquired asymmetry index agrees, regardless of magnitude, so a clinically insignificant asymmetry of -0.02 and a striking -0.50 both score as a match. In Eq. (16), CMAE multiplies each ROI error by AreaROI/Areamax, down-weighting small ROIs such as HipAmy, which is often the epileptogenic focus in temporal lobe epilepsy. The whole-brain delta-SUVR metrics average over all voxels, diluting a focal temporal hypometabolism. The paper's own Table 2 shows the fragility of the 'best model' claim: SGM-KD wins on delta-SUVR/ICC, while SGM-VP wins on CI/CMAE. With only 10 test subjects and no reader study (explicitly deferred in Section 5), the clinical accuracy claim is not supported.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper compares three deep learning models—SGM-KD, SGM-VP, and TransUnet—for synthesizing full-dose FDG brain PET from MRI alone or from MRI plus 1% ultralow-dose PET, using data from 52 subjects (40 train, 2 validation, 10 test) scanned with simultaneous PET/MRI. The authors report voxel-wise SUVR error, ICC, and two newly proposed epilepsy-specific metrics (Congruence Index and Congruence Mean Absolute Error) that quantify hemispheric asymmetry agreement. They conclude that diffusion models show strong potential for pure MRI-to-PET translation and that all three model types can synthesize full-dose FDG-PET accurately when MRI and ultralow-dose PET are available.","tokens_in":15002,"tokens_out":2815,"duration_ms":32942,"significance":"The study addresses a clinically relevant problem—reducing or eliminating radiation exposure in epilepsy FDG-PET—and has notable strengths: the use of list-mode data to simulate realistic 1% ultralow-dose PET, the inclusion of epilepsy-specific hemispheric-asymmetry metrics rather than generic image similarity metrics alone, and a head-to-head comparison of two score-based diffusion models with a transformer-based U-Net. If the quantitative claims were fully supported, the paper would provide a useful benchmark for MRI-to-PET synthesis in epilepsy. However, the central conclusion that the models synthesize full-dose PET 'accurately' in a clinical sense is not established by the presented evidence, because no reader study or diagnostic-endpoint evaluation is performed, the test set is small, and the proposed metrics remain unvalidated against clinically meaningful outcomes. The paper itself acknowledges in Section 5 that assessing whether the Congruence Index correlates with reader studies is future work, so the abstract and conclusion currently overstate the strength of the findings.","major_comments":[{"comment":"The claim that 'all 3 model types can synthesize full-dose FDG-PET accurately' is not supported by the evidence presented. Section 5 explicitly defers assessment of whether the Congruence Index 'correlates with findings from reader studies' to future work, and no reader study or diagnostic-localization endpoint is reported. Equations (15) and (16) quantify sign agreement and area-weighted error of hemispheric asymmetry, but they cannot establish that the synthesis preserves clinically meaningful focal hypometabolism: Eq. (15) counts an ROI as congruent whenever the sign of the asymmetry index agrees, regardless of magnitude, and Eq. (16) down-weights small ROIs such as HipAmy, which is often the epileptogenic focus in temporal lobe epilepsy. The conclusion should be tempered to state that the synthetic images match the acquired PET on quantitative similarity metrics, and that clinical utility remains to be demonstrated.","section":"Abstract and Conclusion (Section 5)"},{"comment":"The 'best model' claims are not statistically robust. In Table 2, SGM-KD with T1w+T2-FLAIR has the highest SUVR ICC (0.84, 95% CI 0.70–0.92) and the lowest delta-SUVR mean in Table 4 (0.96, 95% CI 0.81–1.10), but these confidence intervals overlap substantially with those of several other model-input combinations, including SGM-KD with T1w alone (ICC 0.82, CI 0.67–0.91; delta-SUVR 1.04, CI 0.88–1.20) and TransUnet with T1w alone (delta-SUVR 1.16, CI 0.95–1.37). No formal pairwise testing or correction for multiple comparisons is reported, so the designation of a single best model is not justified by the data. The authors should either add appropriate statistical tests or soften the ranking language throughout the abstract and results.","section":"Tables 2 and 4"},{"comment":"The statement that 'all models improve significantly' when 1% PET input is added is not supported by any significance test. The paper reports confidence intervals for CI, CMAE, and delta-SUVR, but it does not report paired comparisons between the zero-dose and ultralow-dose conditions, nor p-values or effect sizes for any metric. Given the overlapping intervals in Tables 2–4, the word 'significantly' should be removed or substantiated with an appropriate paired statistical analysis.","section":"Abstract and Results (Tables 2 and 3)"},{"comment":"The test set consists of only 10 subjects, and the confidence intervals for the primary metrics are correspondingly wide. For example, the ICC intervals in Tables 2 and 3 span ranges such as 0.63–0.90 and 0.59–0.89, and the CI intervals for the zero-dose task include values as low as 0.51–0.71. With this sample size, the paper should avoid generalizing beyond the studied cohort and should explicitly discuss the uncertainty in the model rankings. A per-subject or per-ROI analysis, or bootstrap resampling, would help quantify the stability of the reported rankings.","section":"Section 2.4 and Tables 2–4"}],"minor_comments":[{"comment":"There are typographical and grammatical errors, including 'ultra-lowdose' (should be 'ultralow-dose') and 'SGMs holds' (should be 'SGMs hold').","section":"Abstract"},{"comment":"The inequality notation in the VPSDE description appears corrupted in the manuscript (e.g., '0 ¡ β1 ¡ β2 ¡ .... ¡ βT ¡ 1'), and several equations contain OCR-like artifacts. The authors should ensure that the mathematical notation renders correctly.","section":"Section 2.2"},{"comment":"The table headers contain duplicated column names, such as 'SUVR ICC' appearing twice and 'Congruence Index' twice. The intended layout should be clarified, and unit labels (e.g., 'CMAE (x10^3)') should be placed consistently.","section":"Tables 2 and 3"},{"comment":"The definition of Areamax in Eq. (16) is given in passing as 'usually the cerebral white matter'; this should be stated precisely and justified, since it directly affects the CMAE values and therefore the model rankings.","section":"Section 2.6"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is within the scope of the journal and does not exhibit problematic citation practices. The central technical contribution—a comparison of diffusion and transformer models for epilepsy-specific MRI-to-PET synthesis with list-mode-simulated low-dose inputs—is potentially valuable. The main reason for major revision rather than rejection is that the quantitative results, while preliminary, are worth reporting after the claims are aligned with the evidence. The authors should add appropriate statistical comparisons, temper the 'accurate synthesis' wording, and clearly separate image-similarity performance from clinical utility."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"I'll be direct: this is a worthwhile feasibility study, but the abstract oversells the clinical accuracy. The genuinely new piece is applying score-based diffusion (SGM-KD and SGM-VP) to MRI-to-PET synthesis specifically for epilepsy, including 1% ultra-low-dose list-mode simulation as a conditioning input, and the proposed CI/CMAE metrics for hemispheric asymmetry. The methods are standard and clearly described; the data handling (40 train/2 val/10 test) is honest; and the authors explicitly acknowledge the lack of a reader study and the single-center data. That transparency is to their credit.\n\nThe weak spot is the inference from image-similarity metrics to the claim that all three models 'synthesize full-dose FDG-PET accurately.' The Congruence Index in Eq. (15) only checks sign agreement, so a clinically meaningless asymmetry of -0.02 and a striking -0.50 both score as a match. The CMAE in Eq. (16) weights errors by ROI area, down-weighting the small hippocampus/amygdala region that is often the epileptogenic focus. Whole-brain delta-SUVR averages over all voxels and can dilute focal temporal hypometabolism. With only 10 test subjects and no reader study, the evidence supports 'promising and needs validation,' not 'accurate.' The paper's own conclusion admits this future work, so the fix is mostly to temper the abstract.\n\nAlso, the model ranking is metric-dependent: SGM-KD wins on delta-SUVR/ICC, while SGM-VP wins on CI/CMAE. That fragility is acknowledged in the discussion, but it underscores that the differences are small and unlikely to be robust with n=10.\n\nBottom line: this is a solid feasibility study with honest limitations. It deserves a serious referee, though the referee should push for softer claims, a reader study (or clear future-work labeling), and per-ROI reporting for small structures. I'd accept it for peer review and would cite it as a comparison if I were working on PET synthesis.","headline":"A useful feasibility study with an overreaching abstract; the clinical accuracy claim needs a reader study and better metrics, but the work deserves review.","tokens_in":15524,"tokens_out":2614,"would_cite":true,"duration_ms":25505,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Score-based diffusion models can synthesize full-dose FDG-PET from MRI alone with metabolic accuracy close to acquired scans in epilepsy patients.","keywords":["score-based generative models","diffusion models","MRI-to-PET translation","epilepsy","FDG-PET","image synthesis","SUVR","ultralow-dose PET"],"falsifier":"A blinded reader study in which epilepsy specialists mark the suspected focus on synthetic zero-dose PET and on the acquired full-dose PET would settle the clinical claim: if localization accuracy on the synthetic images does not agree with the acquired scan beyond chance, or if the Congruence Index disagrees with expert lateralization on the same cases, the central claim would be falsified.","tokens_in":14548,"feed_emoji":"🧠","tokens_out":8227,"duration_ms":76311,"temperature":0.7,"pith_summary":"The paper asks whether deep learning can generate diagnostically usable full-dose FDG (fluorodeoxyglucose) brain PET images without any PET radiation, or with only 1% of the usual dose, for patients being evaluated for epilepsy. Using simultaneous PET/MRI from 52 subjects, it compares two score-based generative diffusion models (SGM-KD and SGM-VP) with a Transformer-U-Net baseline for three input combinations: T1-weighted MRI alone, T1w plus T2-FLAIR, and those MRIs plus 1% ultralow-dose PET. The authors report that SGM-KD synthesizing purely from T1w and T2-FLAIR achieves the lowest whole-brain SUVR error and highest SUVR reliability among the zero-dose models, and that adding 1% PET input improves all models to the point that they become interchangeable. The motivation is clinical: this patient population is young, and eliminating or reducing radiation while preserving metabolic quantitation would be a direct benefit.","feed_headline":"MRI-only PET synthesis nears full-dose FDG accuracy in epilepsy","feed_subtitle":"A diffusion model matched metabolic quantitation targets with zero radiation; adding just 1% of the PET dose made all models equivalent.","key_machinery":"The load-bearing mechanism is conditional score-based diffusion: the forward process gradually corrupts the full-dose PET volume to Gaussian noise following a stochastic differential equation, and a neural network learns the score (the gradient of the log data density) conditioned on the MRI and/or 1% PET inputs; sampling reverses the SDE to generate the PET. Two variants are used: SGM-VP with a variance-preserving SDE and predictor-corrector sampling, and SGM-KD with the denoiser $D_\\theta(x,\\sigma,y)=c_\\text{skip}(\\sigma)x+c_\\text{out}(\\sigma)F_\\theta(c_\\text{in}(\\sigma)x;c_\\text{noise}(\\sigma),y)$ and the Karras stochastic sampler. The clinical evaluation is carried by the Congruence Index and Congruency Mean Absolute Error, which score agreement of left-right SUVR asymmetry across eight paired regions of interest.","core_discovery":"On the paper's own terms, the central discovery is that conditional score-based diffusion models can synthesize full-dose FDG-PET from MRI alone accurately enough to reproduce standard metabolic quantification, and that all tested models produce accurate full-dose PET when supplied with T1w, T2-FLAIR, and 1% ultralow-dose PET. In the zero-dose task, SGM-KD with T1w and T2-FLAIR inputs had the best whole-brain voxel-wise SUVR accuracy (mean $\\Delta$SUVR of $0.96\\times10^{-2}$) and the highest SUVR ICC (0.84), while SGM-VP was best at reproducing hemispheric asymmetry as measured by the proposed Congruence Index and CMAE. The paper introduces these congruence metrics specifically because hemispheric metabolic asymmetry is central to epilepsy focus localization, and argues that standard image metrics such as SSIM and PSNR do not capture this. It also reports that including 1% PET turns the task into a denoising problem, with all models reaching Congruence Index values around 0.85-0.90 and TransUnet the fastest to sample.","pith_inferences":["The paper leaves open whether the Congruence Index tracks expert judgment; a natural extension is a reader study comparing focus localization on synthetic versus acquired PET, and if the index predicts expert lateralization, the metrics could become a standard for PET synthesis evaluation.","Because the authors attribute some performance differences to limited dataset size, the ranking between SGM-VP and SGM-KD may shift with larger multi-center training cohorts; the 10-subject test set is too small to settle architecture superiority.","The 2D slice-based approach causes visible slice inconsistencies in coronal and sagittal views, so 3D or slice-consistent diffusion formulations are an obvious next step that could improve the zero-dose task more than adding MRI contrasts.","If the 1% PET result transfers from list-mode undersampling of full-dose data to true low-dose acquisitions, ultralow-dose reconstruction could replace full-dose imaging in routine PET/MRI epilepsy protocols."],"forward_implications":["If the zero-dose SGM-KD result holds, FDG-PET-like metabolic quantification could be obtained from MRI alone, removing radiation exposure for epilepsy workup.","The finding that all models become interchangeable with 1% PET input suggests a 99% dose reduction could be clinically feasible with any of the three architectures.","Score-based diffusion models are the stronger choice for MRI-only synthesis, while the faster TransUnet becomes competitive once ultralow-dose PET is available.","The proposed Congruence Index and CMAE give a quantitative target for whether a synthetic PET preserves the hemispheric asymmetry pattern that matters for epilepsy diagnosis."],"supporting_citations":[{"why":"Supplies the score-based SDE framework, backward SDE, and predictor-corrector sampler used by the SGM-VP model.","marker":"[15]"},{"why":"Supplies the SGM-KD denoiser formulation and stochastic sampler used for the best zero-dose model.","marker":"[16]"},{"why":"Supplies the VPSDE formulation and PET reconstruction application that SGM-VP is adapted from.","marker":"[13]"},{"why":"Supplies the TransUnet architecture and the prior MRI-to-PET prediction baseline this study compares against.","marker":"[9]"},{"why":"Supplies the list-mode random event-undersampling method used to simulate 1% ultralow-dose PET from full-dose data.","marker":"[17]"},{"why":"Supplies the Asymmetry Index formulation that the proposed Congruence Index and CMAE are built from.","marker":"[22]"},{"why":"Documents altered hemispheric symmetry in temporal lobe epilepsy, motivating the congruence metrics.","marker":"[19]"},{"why":"Supports the claim that simulated ultralow-dose PET approximates true low-dose acquisitions.","marker":"[35]"}],"fun_headline_variants":["Diffusion model synthesizes full-dose PET from MRI alone","Zero-dose PET? Diffusion generates full-dose images from MRI","MRI-only PET synthesis matched full-dose metabolic accuracy","Diffusion model best for MRI-only PET; 1% dose equalizes all models","New congruence metrics capture epilepsy-relevant PET asymmetry"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The claim that the synthetic PET images are accurate for epilepsy rests on the assumption that the Congruence Index, CMAE, and SUVR-based reliability metrics capture what clinicians need to localize epileptogenic zones, an assumption the paper explicitly leaves unvalidated because no reader study was performed.","fun_headline_variants_meta":{"raw":{"variants":["Diffusion model synthesizes full-dose PET from MRI alone","Zero-dose PET? Diffusion generates full-dose images from MRI","MRI-only PET synthesis matched full-dose metabolic accuracy","Diffusion model best for MRI-only PET; 1% dose equalizes all models","New congruence metrics capture epilepsy-relevant PET asymmetry"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000638,"raw_usage":{"total_tokens":3029,"prompt_tokens":1128,"completion_tokens":1901,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":744,"completion_tokens_details":{"reasoning_tokens":1818}},"tokens_in":744,"tokens_out":1901,"duration_ms":16811,"temperature":1.0,"reasoning_tokens":1818,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T04:11:41.177389+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A blinded reader study in which epilepsy specialists mark the suspected focus on synthetic zero-dose PET and on the acquired full-dose PET would settle the clinical claim: if localization accuracy on the synthetic images does not agree with the acquired scan beyond chance, or if the Congruence Index disagrees with expert lateralization on the same cases, the central claim would be falsified.","supporting_citations":[{"cited_title":"Predicting FDG-PET Images From Multi-Contrast MRI Using Deep Learning in Patients With Brain Neoplasms","cited_arxiv_id":null,"evidence_quote":"Supplies the TransUnet architecture and the prior MRI-to-PET prediction baseline this study compares against."},{"cited_title":"Ultra–Low-Dose 18F-Florbetaben Amyloid PET Imaging Using Deep Learning with Multi-Contrast MRI Inputs","cited_arxiv_id":null,"evidence_quote":"Supplies the list-mode random event-undersampling method used to simulate 1% ultralow-dose PET from full-dose data."},{"cited_title":"Enhancing the Diagnostic Utility of ASL Imaging in Tempo- ral Lobe Epilepsy through FlowGAN: An ASL to PET Image Translation Framework","cited_arxiv_id":null,"evidence_quote":"Supplies the Asymmetry Index formulation that the proposed Congruence Index and CMAE are built from."},{"cited_title":"Altered hemispheric symmetry found in left-sided mesial temporal lobe epilepsy with hippocampal sclerosis (MTLE/HS) but not found in right-sided MTLE/HS","cited_arxiv_id":null,"evidence_quote":"Documents altered hemispheric symmetry in temporal lobe epilepsy, motivating the congruence metrics."},{"cited_title":"True ultra-low-dose amyloid PET/MRI enhanced with deep learning for clinical interpretation","cited_arxiv_id":null,"evidence_quote":"Supports the claim that simulated ultralow-dose PET approximates true low-dose acquisitions."}],"review_version":1}