{"id":"0f30e414-b488-4155-84fa-a5c8a03f285e","arxiv_id":"2411.09512","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":0.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A review of GAN-based low-dose CT denoising methods that presents no original results.","lead":"This paper reviews how generative adversarial networks, or GANs, can clean up noisy low-dose CT scans. It surveys many published methods, but contains no new experiments or results.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Section 4.2's comparison table is the quantitative backbone of the survey, and it is not trustworthy: metrics are drawn from incompatible protocols, and the SRGAN row appears to report a non-GAN super-resolution network as GAN evidence.","rationale":"The reader correctly identified the incommensurability of the metrics in Section 4.2 as a load-bearing weakness. My stress-test agrees with that reading but sharpens it: the table is not merely hard to compare across rows; at least one row appears to misclassify a non-GAN method as SRGAN, which would directly contaminate the evidence for the central claim. I do not think this is a fatal objection to the paper's genre or topic, because the broader literature does contain GAN methods that report improvements on LDCT denoising. The paper can be repaired by rebuilding the table with strict per-row sourcing, matching dataset and metric definitions, and removing or relabeling any non-GAN methods. That is a conditional acceptance path, not a rejection of the reviewed literature. The verdict remains CONDITIONAL as the reader recommended, so no change in verdict is needed.","tokens_in":12700,"tokens_out":4583,"duration_ms":45049,"concrete_test":"Trace the SRGAN row of Table 4.2 to Chi et al. [11]: retrieve the source paper and determine whether the model is trained with a generator and discriminator in an adversarial objective, and record the exact meaning of each reported number (metric, dataset, scaling factor, dose level). If the network is not a GAN, or if the reported values are PSNR figures placed in the SSIM column, the table cannot serve as evidence for GAN-based superiority.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that GAN architectures 'revolutionize' low-dose CT denoising is supported mainly by the assembled quantitative results in Section 4.2. That table cannot bear the weight. It mixes SSIM, PSNR, LPIPS, GMSD, and RMSE values from LIDC-IDRI, 3D-IRCADB, PANCREAS, Mayo Clinic, and phantom studies that use different dose levels, noise simulations, reconstruction pipelines, and evaluation protocols, with no normalization or per-row sourcing. Several rows are internally misaligned: the CycleGAN row reports only '~44' with no metric label, and the DU-GAN row lists 'ss 23.1102 0.0724' without an SSIM value. More seriously, the SRGAN row reports numbers attributed to Chi et al. [11], but the architecture described in Section 3.3 (GDAFM, MAB, FFDM) contains no generator-discriminator game or adversarial loss; it is presented as a super-resolution network, not a GAN. If [11] is not a GAN, then the table includes a non-GAN method as direct evidence for GAN superiority. The paper itself concedes in Section 5.1 that PSNR and SSIM 'do not always represent clinically useful images,' yet the conclusion is clinically framed. The underlying claim could still be true in the broader literature, but this review's quantitative synthesis does not establish it.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript is a narrative review of generative adversarial network (GAN) architectures applied to low-dose computed tomography (LDCT) denoising. It covers conditional GANs, CycleGANs, SRGANs, denoising GANs, DualGAN/DU-GAN, and Wasserstein GANs, and it summarizes reported performance on datasets such as LIDC-IDRI, 3D-IRCADB, PANCREAS, Mayo Clinic, and phantoms. The paper closes with technical and clinical challenges and proposes future directions. The abstract and conclusion assert that GAN-based methods are 'revolutionary' and provide 'an advanced resolution' to the enduring trade-off between radiation exposure and image quality.","tokens_in":12959,"tokens_out":5110,"duration_ms":43962,"significance":"The paper gathers a broad set of recent references and describes several GAN variants, which could be useful as a qualitative entry point for readers new to LDCT denoising. It also explicitly acknowledges important limitations, including synthetic artifacts, interpretability, and the limited clinical relevance of PSNR/SSIM. However, the quantitative evidence in Section 4.2 is too inconsistent and malformed to support the strong evaluative claim. One table entry is based on a non-GAN method, and the paper itself concedes that standard metrics do not necessarily reflect clinical utility. The review is therefore better characterized as a qualitative scoping overview than as a rigorous critical synthesis that substantiates the 'revolutionizing' narrative.","major_comments":[{"comment":"The comparison table is the quantitative backbone of the review, but it is not usable in its current form. It mixes SSIM, PSNR, LPIPS, GMSD, and RMSE values drawn from LIDC-IDRI, 3D-IRCADB, PANCREAS, Mayo Clinic, and phantom studies that use different dose levels, noise simulations, reconstruction pipelines, and evaluation protocols, with no normalization or per-row specification of the evaluation conditions. Several entries are uninterpretable: the CycleGAN row reports only '~44' with no metric label, the DU-GAN row reports 'ss 23.1102 0.0724' without a separate SSIM value, and the SRGAN row lists two SSIM/PSNR pairs without indicating which dataset each corresponds to. No error bars, confidence intervals, or significance tests are provided. This table therefore cannot support the conclusion that GAN variants outperform one another or that GANs are superior to non-GAN methods.","section":"Section 4.2, Table 'Results Comparison between GAN variants'"},{"comment":"The SRGAN section conflates a genuine GAN, that of Ledig et al. [10], with the non-adversarial super-resolution architecture of Chi et al. [11]. The description of Chi et al.'s method includes GDAFM, MAB, and FFDM but never mentions a generator-discriminator game or an adversarial loss; it is presented as a super-resolution reconstruction network. Nevertheless, Section 4.2 attributes the Chi et al. numbers to the 'SRGAN' row, using a non-GAN method as evidence for GAN performance. This is a load-bearing attribution error that directly affects the quantitative claim; the row should be removed or clearly labeled as a non-GAN baseline.","section":"Section 3.3 and Section 4.2"},{"comment":"The paper acknowledges in Section 5.1 that 'High numerical scores do not always represent clinically useful images,' yet the abstract and conclusion rely on quantitative PSNR/SSIM improvements as reassurance of the 'transformative' and 'revolutionizing' potential of GANs for clinical practice. The review does not provide a clinically validated evaluation framework, nor does it reconcile the acknowledged metric-reality gap with the strength of its conclusion. As written, the central claim is not commensurate with the evidence presented; a revision should either temper the conclusion or supply a clinically oriented assessment that makes the quantitative results meaningful.","section":"Section 5.1 and Conclusion"}],"minor_comments":[{"comment":"The manuscript contains many grammatical errors and garbled sentences, such as 'deoxidized pictures' in the Figure 1 caption and 'the network will certainly modify the input data' in Section 5.2. A thorough language edit is needed.","section":"Throughout"},{"comment":"The text states 'Zhang et al. first proposed this concept in their paper [12]', but reference [12] is authored by Chen et al. This author-name mismatch should be corrected.","section":"Section 3.4"},{"comment":"Reference [17] is cited to support the Wasserstein GAN equations, but [17] is a paper on DCGAN-based defect detection, not a primary or authoritative source for WGAN mathematics; a more appropriate citation (e.g., Arjovsky et al. [16]) should be used.","section":"Section 3.6"},{"comment":"The RMSE formula has a typographical issue: the text says 'the i-th observed value and i-th predicted value' but the formula is missing the subscripted hat for the predicted value. This should be corrected to match the standard definition.","section":"Section 4.1.4, Eq. (5)"},{"comment":"The table is not numbered or given a caption in the text, making it difficult to refer to. Please add a proper table number and a caption that describes the sources and the exact metric definitions used in each row.","section":"Section 4.2"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a literature review without original experiments, and the fit with a venue expecting technical contributions may be questionable. More importantly, the quantitative synthesis is currently unreliable; the table in Section 4.2 cannot support the paper's strong claims. I would advise the editor that a major revision is warranted, and that the authors should either substantially repair the table and remove non-GAN entries or restructure the paper as a purely qualitative review without the quantitative 'comparison' framing."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This is a review paper, not original research, and the one quantitative table that is supposed to be its backbone is not trustworthy. The prose walk through cGANs, CycleGANs, WGANs, and the rest is a reasonable reading list for someone new to the area, and it does flag real issues like synthetic artifacts, interpretability, and the gap between PSNR/SSIM and clinical usefulness. The reference list is broad and mostly standard. Credit where it's due: the survey covers the major architecture families and honestly acknowledges that high metric scores don't necessarily mean good diagnostic images.\n\nThe soft spots are serious, though. Section 4.2's comparison table mixes PSNR, SSIM, LPIPS, GMSD, and RMSE values from different datasets, dose levels, and evaluation protocols with no normalization and no per-row sourcing. The CycleGAN row lists only '~44' with no metric label; the DU-GAN row reads 'ss 23.1102 0.0724' with no SSIM value. And the SRGAN row cites Chi et al. [11], which the paper itself describes as a super-resolution network with a feedback feature distillation mechanism, not a GAN. That is not adversarial training. So the table includes a non-GAN method as direct evidence for GAN superiority. The stress-test note is right: the quantitative synthesis cannot bear the weight of the abstract's 'revolutionizing' claim.\n\nThe paper does concede in Section 5.1 that PSNR and SSIM 'do not always represent clinically useful images,' but that concession is exactly why the table is a problem. If the numbers can't be compared, the central survey claim loses its quantitative support. The writing is also rough in places—typos, garbled sentences, and at least one 'audio removal' slip—which makes the whole thing feel unpolished.\n\nWho gets value from this? A graduate student wanting a quick map of the GAN-for-LDCT literature might use the reference list. Anyone needing a reliable quantitative summary should look elsewhere. My recommendation: a serious editor would want this heavily revised before sending it to referees. It is not ready as is. The authors need to discipline the survey with a stated selection method, fix the table with standardized metrics and clear sourcing, and dial back the language. I would not cite it and I would not put it on a reading list except as a cautionary example of how not to run a comparative review.","headline":"A broad but sloppy review of GAN-based low-dose CT denoising; the prose survey has some use, but the quantitative comparison table is broken and cannot support the paper's central claims.","tokens_in":13443,"tokens_out":2350,"would_cite":false,"duration_ms":23065,"reading_group":"no","serious_thinker":"no","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This review argues that GAN-based architectures have become an effective solution to the low-dose CT trade-off between radiation exposure and image quality.","keywords":["low-dose computed tomography","denoising","generative adversarial networks","conditional GAN","CycleGAN","Wasserstein GAN","image quality metrics","medical imaging"],"falsifier":"A standardized re-evaluation would settle the claim: take the cited GAN variants (cGAN, CycleGAN, SRGAN, WGAN, DU-GAN, PWGAN) and train each on the same low-dose CT dataset, e.g., the Mayo Clinic AAPM Low Dose CT Grand Challenge, with identical training and validation splits and preprocessing, then compare PSNR, SSIM, and LPIPS on a common test set. If the reported performance rankings reverse or the claimed margins shrink to noise under this controlled comparison, the survey's central quantitative claim would not survive.","tokens_in":12496,"feed_emoji":"🩻","tokens_out":4089,"duration_ms":33620,"temperature":0.7,"pith_summary":"This review argues that GAN-based architectures have matured into a practical answer to the central dilemma of low-dose CT: reducing radiation exposure without sacrificing diagnostic image quality. By surveying cGANs, CycleGANs, SRGANs, denoising GANs, DualGANs, and Wasserstein GANs, the paper claims that adversarial training lets generators map noisy low-dose images to normal-dose-like outputs while preserving anatomical detail. The review reports quantitative gains in PSNR, SSIM, and LPIPS across benchmark and clinical datasets and concludes that GAN-based denoising holds promise for precision medicine, provided remaining barriers are addressed.","feed_headline":"GAN-based denoising closes CT's dose-versus-quality gap","feed_subtitle":"A review of cGAN, CycleGAN, SRGAN and WGAN variants shows gains in PSNR and SSIM, with clinical hurdles remaining.","key_machinery":"The load-bearing mechanism is the adversarial generator–discriminator pair, in which a generator learns to produce denoised images and a discriminator learns to distinguish them from real normal-dose CT images. The review tracks how each architecture modifies this core: cGANs condition generation on the input image, CycleGANs add cycle-consistency losses to work with unpaired data, SRGANs add perceptual losses for super-resolution, and WGAN variants replace the Jensen–Shannon objective with the Wasserstein distance for training stability. Across all variants, the decisive design choices are the loss-function combination and the discriminator's domain (image, gradient, or both), which together control the balance between noise suppression and artifact introduction.","core_discovery":"The paper's central claim is that GAN-based denoising has moved from a theoretical possibility to a demonstrated technique: across cGAN, CycleGAN, SRGAN, denoising GAN, DualGAN, and WGAN variants, adversarial training lets a generator map low-dose CT images to images that match normal-dose quality in PSNR, SSIM, and perceptual similarity, while preserving anatomical detail. The review also claims that no single architecture wins outright; the gains come from combining adversarial objectives with auxiliary losses—cycle consistency for unpaired data, sharpness or structural losses, and dual-domain discriminators that police both pixels and edges. The paper's conclusion is that these methods hold promise for precision medicine through personalized denoising models, provided current barriers around synthetic artifacts, interpretability, and clinically meaningful evaluation are addressed.","pith_inferences":["The paper's own admission that PSNR and SSIM do not track clinical utility suggests the field's next bottleneck is task-based evaluation, such as radiologist reader studies or lesion-detection sensitivity, rather than more architecture variants.","If the comparative table is not commensurable, then a public benchmark with fixed training data and unified metrics would be a low-cost way to make the next generation of claims testable.","One testable extension: inject controlled noise into a common phantom dataset to isolate architectural improvements from dataset effects, which the review's mixed evidence cannot currently separate.","The review's emphasis on unpaired training points toward leveraging large archives of routine clinical CT scans as unlabeled training data, which could address the paired-data scarcity it identifies."],"forward_implications":["Unpaired image translation via cycle consistency removes the requirement for perfectly aligned low-dose and normal-dose CT pairs, easing a major clinical data bottleneck.","Hybrid loss functions that combine adversarial feedback with perceptual, sharpness, or structural similarity losses mitigate the oversmoothing typical of pure MSE training.","Wasserstein-based objectives stabilize GAN training and are reported to improve artifact removal and detail preservation in dental, lung, and liver CT.","Dual-domain discriminators that evaluate both image pixels and gradients improve edge preservation and reduce streak artifacts.","Clinical adoption remains limited by synthetic artifacts, poor interpretability, computational cost, and the mismatch between numerical metrics and diagnostic usefulness."],"supporting_citations":[{"why":"Supplies the adversarial training objective that all reviewed architectures build on.","marker":"[1]"},{"why":"Provides the early demonstration of GAN-based noise reduction in low-dose CT with adversarial plus voxel-wise loss.","marker":"[4]"},{"why":"Shows a cGAN with U-Net generator for low-dose chest CT denoising and reports SSIM gains.","marker":"[7]"},{"why":"Introduces a sharpness-aware cGAN that preserves edges while denoising paired upper-body CT scans.","marker":"[8]"},{"why":"Defines CycleGAN and the cycle-consistency mechanism enabling unpaired image translation.","marker":"[9]"},{"why":"Establishes SRGAN and perceptual loss for super-resolution, later adapted to LDCT super-resolution and denoising.","marker":"[10]"},{"why":"Presents DU-GAN with dual-domain (image and gradient) discriminators for edge-preserving denoising.","marker":"[15]"},{"why":"Introduces Wasserstein GAN and the Wasserstein distance for stable adversarial training.","marker":"[16]"},{"why":"Reports a progressive WGAN with structure-sensitive hybrid loss that outperforms existing methods on low-dose CT.","marker":"[22]"},{"why":"Defines LPIPS, the perceptual metric the review uses alongside PSNR and SSIM.","marker":"[25]"}],"fun_headline_variants":["GANs lift low-dose CT to full-dose quality, review finds","Adversarial losses, not architecture, drive low-dose CT denoising","cGAN, CycleGAN, SRGAN: six GAN types tested for CT denoising","GAN denoising for CT: PSNR gains, but clinical use lags"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The claim that GAN variants outperform one another depends on treating the PSNR, SSIM, and LPIPS values collected from different studies, datasets, and evaluation protocols as directly comparable evidence; the paper itself acknowledges that those numeric scores do not always correspond to clinically useful images.","fun_headline_variants_meta":{"raw":{"variants":["GANs lift low-dose CT to full-dose quality, review finds","Adversarial losses, not architecture, drive low-dose CT denoising","cGAN, CycleGAN, SRGAN: six GAN types tested for CT denoising","GAN denoising for CT: PSNR gains, but clinical use lags"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00091,"raw_usage":{"total_tokens":4167,"prompt_tokens":953,"completion_tokens":3214,"prompt_tokens_details":{"cached_tokens":896},"prompt_cache_hit_tokens":896,"prompt_cache_miss_tokens":57,"completion_tokens_details":{"reasoning_tokens":3127}},"tokens_in":57,"tokens_out":3214,"duration_ms":279368,"temperature":1.0,"reasoning_tokens":3127,"cache_read_input_tokens":896,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T20:34:02.049172+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A standardized re-evaluation would settle the claim: take the cited GAN variants (cGAN, CycleGAN, SRGAN, WGAN, DU-GAN, PWGAN) and train each on the same low-dose CT dataset, e.g., the Mayo Clinic AAPM Low Dose CT Grand Challenge, with identical training and validation splits and preprocessing, then compare PSNR, SSIM, and LPIPS on a common test set. If the reported performance rankings reverse or the claimed margins shrink to noise under this controlled comparison, the survey's central quantitative claim would not survive.","supporting_citations":[{"cited_title":"& Bengio, Y.(2020).Generativeadversarialnetworks.CommunicationsoftheACM,63(11),139-144","cited_arxiv_id":null,"evidence_quote":"Supplies the adversarial training objective that all reviewed architectures build on."},{"cited_title":"M., Leiner, T., Viergever, M","cited_arxiv_id":null,"evidence_quote":"Provides the early demonstration of GAN-based noise reduction in low-dose CT with adversarial plus voxel-wise loss."},{"cited_title":"J., & Lee, D","cited_arxiv_id":null,"evidence_quote":"Shows a cGAN with U-Net generator for low-dose chest CT denoising and reports SSIM gains."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces a sharpness-aware cGAN that preserves edges while denoising paired upper-body CT scans."},{"cited_title":"[10]Ledig, C.,Theis, L., Huszár, F., Caballero,J., Cunningham,A.,Acosta, A.,","cited_arxiv_id":null,"evidence_quote":"Defines CycleGAN and the cycle-consistency mechanism enabling unpaired image translation."}],"review_version":1}