{"id":"071f6ea7-eecd-434b-8853-d028a9d5504f","arxiv_id":"2508.00438","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A user-guided diffusion inpainting pipeline synthesizes coronary angiograms with controlled stenosis severity, improving downstream lesion detection and severity classification over training on real data alone.","lead":"The paper proposes a diffusion-based data augmentation method that generates synthetic coronary angiograms with user-specified stenosis severity, then uses them to train a YOLO detector. It reports improved lesion detection and severity classification on internal and public datasets, especially when real training data are scarce.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The synthetic images' target %DS is never verified; if the diffusion inpainting does not faithfully render the modified masks, the synthetic labels are wrong and the reported gains may be spurious.","rationale":"I agree with the reader's weakest assumption. The paper's mechanism is only as sound as the mapping from a user-specified %DS to the rendered image, and that mapping is never quantitatively validated. This is the single most load-bearing concern because it affects both internal and external results and the basic claim of 'user-guided control of severity.' The external re-annotation issue (§3.1) is a related but secondary threat to the ARCADE evaluation, and the absence of error bars and simple augmentation baselines weakens the strength of the empirical comparison, but neither is as foundational as the correctness of the synthetic labels themselves. If the proposed QCA re-measurement shows that generated images match their targets, the central mechanism is supported; if not, the reported improvements are very hard to interpret. The reader already assigned CONDITIONAL, and this concern justifies that judgment rather than moving it, so the verdict should remain unchanged pending the proposed validation.","tokens_in":7453,"tokens_out":5941,"duration_ms":60899,"concrete_test":"Re-run the same QCA tool used for annotation (§2.1) on a random sample of, say, 100 generated images from the ×4 balanced set, stratified 50 moderate-target and 50 severe-target, and compare the measured %DS to the target %DS used as the training label. Compute the mean absolute error and the confusion rate at the 70%DS boundary. If the measured severity disagrees with the target label in more than ~10% of cases, or if severe-target images systematically measure below 70%DS, the synthetic labels are unreliable and the reported performance gains in Table 2 may be spurious. This check is directly feasible because the QCA tool and generated images are both available to the authors.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim requires that a synthetic image labeled with a target %DS actually exhibits that stenosis severity, because the downstream YOLO detector is trained to associate image content with these labels (§2.3, §3.3). The authors modify vessel contours to a target %DS (§2.1), condition a ControlNet on the modified mask (§2.2), and then use the target class as ground truth for training. Nowhere do they re-apply QCA to a generated image or ask clinicians to grade a generated image to confirm that the rendered stenosis matches the target; Fig. 1 shows only selected qualitative examples. If the inpainting does not faithfully realize the narrowed mask (e.g., because the diffusion model smooths the MLD, alters the vessel, or introduces artifacts), then the synthetic labels are systematically wrong, and the detector learns incorrect associations between appearance and severity. The reported gains are small (internal F1 0.650→0.670, mAP50 0.688→0.717; ARCADE F1 0.500→0.519→0.524→0.513) and are reported without error bars or significance tests, so a moderate label-error rate could plausibly produce or erase these differences. This is the most load-bearing weakness because every downstream result inherits the unverified assumption that the generative model respects the user-specified %DS.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes DiGDA, a diffusion-based data augmentation pipeline for coronary stenosis detection. The method uses a QCA tool to extract vessel contours and lesion locations from real angiograms, lets a user modify the vessel mask to a target percentage diameter stenosis (%DS), and trains a multi-ControlNet diffusion model to inpaint the lesion region so that the generated image reflects the requested severity. Synthetic images are then added to the training set at ×1, ×2, and ×4 ratios with a class-balanced selection, and a YOLO detector is trained for joint lesion detection and severity classification (moderate vs severe). Experiments on a large internal dataset and on the public ARCADE dataset report higher F1 and mAP50 than a real-data-only baseline, along with ablations on class balance and data scarcity.","tokens_in":7674,"tokens_out":5067,"duration_ms":50264,"significance":"If the reported results are reliable, the paper offers a practical solution to two persistent problems in coronary angiography analysis: limited labeled data and class imbalance. The targeted inpainting strategy is a reasonable extension of ControlNet-based medical image synthesis, and the promise to release a re-annotated ARCADE validation set is a useful contribution to the community. The main value would lie in showing that user-controlled severity modification can inject clinically meaningful variation into training data. However, the central claim rests on an unverified assumption about the faithfulness of the generated images to the specified %DS, and the statistical evidence for the quantitative gains is currently thin; the significance of the work therefore depends on whether these limitations can be addressed with additional experiments.","major_comments":[{"comment":"The synthetic labels are assigned from the user-specified target %DS that is applied to the vessel mask, but the paper never verifies that the generated angiogram actually exhibits that stenosis severity. Since the downstream YOLO detector is trained to associate image content with these labels, a systematic mismatch between target and rendered %DS could corrupt the training signal and explain part or all of the reported gains. The authors should re-apply their QCA pipeline to a sample of generated images and report the distribution of target vs measured %DS, or provide a clinician reader study grading the generated severity. Without such verification, the central improvement claim is not fully supported.","section":"§2.2, §2.3, Fig. 1"},{"comment":"All downstream results are averages over three runs with no standard deviations, confidence intervals, or significance tests. The phrase \"significant lesion detection performance gain\" for the internal ×4 setting (mAP50 0.717 vs 0.688) is therefore not statistically substantiated. Moreover, the ARCADE results contradict the stated monotonic improvement: the ×4 setting (F1 0.513, mAP50 0.492) performs worse than ×2 (F1 0.524, mAP50 0.501). The authors should add variance information, perform paired tests or confidence intervals, and explicitly discuss the non-monotonic external behavior.","section":"§3.3, Table 2"},{"comment":"The data-scarcity comparison between \"Real-only\" and \"Ours\" does not control for the total number of training images: the \"Ours\" condition adds synthetic data on top of the real subset, so the model receives more training examples. The observed mAP50 gains could therefore reflect increased data quantity rather than the specific usefulness of the generated synthetic images. To support the claim that the method \"maintains high ... performance even when trained with limited data,\" the authors should include a control condition that matches total data size, for example by training on a larger real subset of equal size or on real data with standard augmentation.","section":"§3.3, Fig. 4(b)"},{"comment":"The external validation on ARCADE is based on a re-annotation of the validation set with the same in-house QCA tool that is used to annotate the internal real data and to define the synthetic stenosis labels. This creates a consistency between training and evaluation that is not an independent clinical gold standard. The paper should explicitly acknowledge that ARCADE performance measures agreement with the QCA-based labeling pipeline, and ideally compare against expert clinician grading on a subset to assess clinical validity.","section":"§3.1, §3.3"}],"minor_comments":[{"comment":"The expectation in the loss is written as \"E[∥ϵ − ϵθ(zt,t,ct,cs,cm)∥2]\" but the sampling notation \"t, ct, cs, cm ∼ N(0,1)\" is confusing because the noise ϵ is the only variable drawn from a standard normal; the subscript structure of the expectation should be clarified.","section":"§2.2, Eq. (1)"},{"comment":"The paper states that average scores over three runs are reported, but no standard deviations or variance information appear in Table 2 or Figure 4; the figures would be more informative with error bars or interval shading.","section":"§3.2"},{"comment":"The claim that this is \"the first to integrate an inpainting strategy with a generative model that synthesizes angiograms for lesion detection\" is strong, and the related-work section covers only a few diffusion-based augmentation papers; a broader search might reveal closer prior work on coronary angiography generation.","section":"§1"},{"comment":"The internal dataset is described only as a \"curated collection\" of angiograms; acquisition protocol details, patient selection criteria, and exclusion criteria would improve reproducibility and allow readers to judge the clinical scope.","section":"§3.1"},{"comment":"The qualitative examples show original and generated images but do not overlay the modified vessel masks or the intended MLD location on the generated images; showing these overlays would help the reader assess the correspondence between the requested severity change and the visible output.","section":"§3.4, Fig. 1"},{"comment":"The text says \"convolution network\" where \"convolutional network\" would be more standard, and there is a typo in the Table 1 header \"V alid\"; these small issues should be corrected.","section":"§2.2"}],"recommendation":"major_revision","confidential_remarks":"The core idea is interesting and the paper is clearly written, but the missing verification of synthetic label fidelity is a load-bearing gap that should be addressed before publication. I would urge the editor to require a QCA re-measurement of generated images (or a clinician reader study) and to require the authors to report variance and significance in their quantitative comparisons. The re-annotated ARCADE evaluation is a positive contribution, but it is not an independent gold standard, so the conclusions should be tempered accordingly."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The core idea here is genuinely new and the paper is honestly written. The combination of user-guided vessel-mask modification with multi-ControlNet inpainting, targeting the lesion region rather than synthesizing whole images, is a reasonable step beyond prior diffusion augmentation work, which mostly targeted classification. The ablations on class balance, data scarcity, and model size are thoughtful, and the plan to release the re-annotated ARCADE data is a credit.\n\nThat said, the empirical support for the central claim is weaker than the prose suggests. The gains are small (internal mAP50 0.688 to 0.717; ARCADE F1 0.500 to 0.524), reported as averages over three runs with no error bars or significance tests, so we do not know whether these differences are real or noise. The ARCADE results are non-monotonic: ×4 underperforms ×2 on both F1 and mAP50, which sits uncomfortably with the claim of consistent superiority.\n\nThe load-bearing weakness is the label-verification problem. The synthetic images are labeled with a target %DS derived from the modified vessel mask, but the paper never checks whether the diffusion model actually renders that severity. Re-running QCA on generated images or asking clinicians to grade a sample would settle this. If the inpainting smooths or distorts the narrowing, the synthetic labels are systematically wrong and the downstream detector learns incorrect associations—which could easily produce or erase the modest gains reported.\n\nThe external validation is also partly circular. The same QCA tool that generates synthetic labels is used to re-annotate the ARCADE validation set. Even with clinician refinement, this aligns the label noise in synthetic training data and external evaluation, potentially biasing the comparison in the method's favor. A simpler baseline—like oversampling the minority class or standard geometric augmentation—would help calibrate whether the gains come from the diffusion synthesis specifically or just from rebalancing and more data.\n\nWho is this for? Researchers working on diffusion-based augmentation for detection in medical imaging, and anyone thinking about synthetic label fidelity. The paper deserves a serious referee: the idea is coherent, the pipeline is reproducible in principle, and the claimed direction is plausible. A referee should ask for quantitative verification of the synthetic severity labels, significance tests, and at least one cheap augmentation baseline. With those additions the claims would be much more convincing.","headline":"Sensible pipeline, honestly written, but the headline gains are small, unverified, and partly circular; worth reviewing carefully rather than rejecting.","tokens_in":8234,"tokens_out":1909,"would_cite":false,"duration_ms":21228,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Adding user-specified synthetic lesions to angiography training data improves coronary stenosis detection and severity classification beyond real-data-only training, on both an in-house and a public dataset.","keywords":["coronary stenosis","coronary angiography","diffusion models","data augmentation","inpainting","lesion detection","severity classification","class imbalance"],"falsifier":"Take a sample of generated images, run the same QCA tool used to create the training labels on them, and compare the measured %DS with the user-specified target; if the measured values are systematically off or broadly scattered, the claim that severity control is the mechanism behind the gains is not supported.","tokens_in":7222,"feed_emoji":"🫀","tokens_out":8606,"duration_ms":75642,"temperature":0.7,"pith_summary":"Accurate detection and severity grading of coronary stenosis from angiography is held back by small, imbalanced datasets and expensive manual labels. This paper proposes DiGDA, a data-augmentation pipeline that uses a diffusion inpainting model, conditioned on user-modified vessel segmentation masks and lesion bounding boxes, to generate realistic synthetic lesions at a user-specified percentage diameter stenosis (%DS). The synthetic images are balanced across moderate and severe classes and added to the training set of a YOLO detection-and-classification model. On both a large in-house dataset and a public external dataset, the augmented model reports higher F1, precision, recall, and mAP50 than real-data-only training, with the largest gains concentrated in the severe class, and it stays ahead under data-scarce training subsets. The claim is that this form of targeted, controllable augmentation replaces some of the need for additional expert annotations.","feed_headline":"Synthetic lesions lift stenosis detection score from 0.688 to 0.717","feed_subtitle":"A diffusion model edits vessel masks to create rare severe cases, cutting reliance on scarce expert labels.","key_machinery":"The load-bearing object is the multi-ControlNet inpainting stage: a latent diffusion model, based on Stable Diffusion, whose denoising network $\\epsilon_\\theta$ is fine-tuned with zero-convolution conditioning layers to take two extra inputs -- an edited vessel segmentation mask $c_s$ and a masked image with a lesion bounding box $c_m$ -- along with the text prompt $c_t$ and timestep $t$. The vessel mask is edited by a QCA-based algorithm that moves two control points at the minimum lumen diameter orthogonally to the vessel direction until the computed %DS, $\\%DS = (1 - MLD/D_{\\mathrm{ref}}) \\times 100$, matches the user's target, with smooth propagation to neighboring contour points. These conditions let the model redraw only the stenosed segment while preserving the surrounding angiographic context. The downstream detector is a YOLO model that simultaneously outputs lesion bounding boxes and a moderate/severe class.","core_discovery":"The central claim is that training a single-stage lesion detector on a mix of real angiograms and diffusion-inpainted synthetic angiograms, generated with user-controlled stenosis severity, improves both localization and severity classification relative to training on real data alone. The paper reports mAP50 increasing from 0.688 with no synthetic data to 0.717 with a ×4 balanced synthetic set on the internal test set, and corresponding F1 and mAP50 improvements on the external public dataset. The method restricts generation to targeted lesion regions by inpainting within bounding boxes, using a multi-ControlNet that takes both the vessel segmentation mask (edited to a target %DS) and the masked original image as conditions. This is presented as the first use of inpainting-based diffusion augmentation for lesion detection in coronary angiography, and the authors argue it avoids the unintended lesion artifacts that full-image synthesis can introduce. The severity control is achieved by a QCA-based algorithm that adjusts the minimum lumen diameter in the vessel contour to match the requested %DS.","pith_inferences":["One implication the paper leaves implicit: if the gains come mainly from adding diverse, balanced examples rather than from severity-accurate labels, then a cheaper generator that ignores target %DS and only adds controlled diversity might capture most of the benefit; a comparison against synthetic images with shuffled severity labels would settle this.","A natural extension is to apply the same mask-edit-and-inpaint recipe to other focal vascular lesions, such as aneurysms, dissections, or calcified plaques, where lesions can appear anywhere along a vessel and labeled examples are scarce.","The method's preservation of background context outside the edited bounding box makes it more suited to clinical trust than full-image synthesis, but that trust would still need a reader study where cardiologists rate generated frames for realism and severity before deployment."],"forward_implications":["At a ×4 synthetic-to-real ratio, the internal mAP50 rises from 0.688 to 0.717, with the severe-lesion class improving from 0.646 to 0.663, so the augmentation most helps the underrepresented class.","The external public dataset shows the same trend (F1 from 0.500 to 0.513 at ×4), suggesting the benefit transfers across imaging protocols and patient populations.","A balanced synthetic set outperforms an imbalanced synthetic set of the same total size across model scales, so the class-balancing step is part of why the method works.","Under data-scarce subsets (as low as 5% of the training data), the augmented model maintains higher mAP50 than real-only training, particularly for severe lesions, indicating the method can stretch limited labeled collections."],"supporting_citations":[{"why":"Supplies the ControlNet conditioning architecture that lets vessel masks and bounding boxes steer synthetic image generation.","marker":"[24]"},{"why":"Provides the latent diffusion base model that the inpainting pipeline fine-tunes.","marker":"[20]"},{"why":"Supplies the external public dataset used to test whether the augmentation benefit generalizes beyond the in-house data.","marker":"[19]"},{"why":"Defines the YOLO single-stage detector used for lesion localization and severity classification.","marker":"[10]"},{"why":"Grounds the QCA-based computation of %DS from reference diameter and minimum lumen diameter used to edit vessel masks.","marker":"[3]"},{"why":"Defines quantitative coronary angiography measurement principles, including reference diameter and minimal lumen diameter estimates.","marker":"[4]"},{"why":"Demonstrates controllable diffusion augmentation for object detection and serves as the prior work the paper contrasts with its targeted inpainting approach.","marker":"[8]"},{"why":"Provides the clinical severity thresholds (≥50% clinically significant, ≥70% severe) that define the classification labels and synthetic balancing targets.","marker":"[16]"}],"fun_headline_variants":["Synthetic lesions from diffusion boost stenosis detection","User-guided diffusion adds rare stenosis cases for AI","Inpainting-based augmentation lifts stenosis detection","Detector gains from diffusion-generated severe lesions","Severity-controlled synthetic lesions improve stenosis scoring"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The augmented images must really look like coronary angiograms with the stenosis severity the user asked for; if the detector learns from images that do not match their labels, the reported gains will not transfer to real clinical data.","fun_headline_variants_meta":{"raw":{"variants":["Synthetic lesions from diffusion boost stenosis detection","User-guided diffusion adds rare stenosis cases for AI","Inpainting-based augmentation lifts stenosis detection","Detector gains from diffusion-generated severe lesions","Severity-controlled synthetic lesions improve stenosis scoring"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000172,"raw_usage":{"total_tokens":1267,"prompt_tokens":928,"completion_tokens":339,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":544,"completion_tokens_details":{"reasoning_tokens":272}},"tokens_in":544,"tokens_out":339,"duration_ms":3914,"temperature":1.0,"reasoning_tokens":272,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T10:07:28.843489+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a sample of generated images, run the same QCA tool used to create the training labels on them, and compare the measured %DS with the user-specified target; if the measured values are systematically off or broadly scattered, the claim that severity control is the mechanism behind the gains is not supported.","supporting_citations":[{"cited_title":"In: Proceedings of the IEEE/CVF International Conference on Computer Vision","cited_arxiv_id":null,"evidence_quote":"Supplies the ControlNet conditioning architecture that lets vessel masks and bounding boxes steer synthetic image generation."},{"cited_title":"Scientific Data11(1), 20 (2024) 2, 6","cited_arxiv_id":null,"evidence_quote":"Supplies the external public dataset used to test whether the augmentation benefit generalizes beyond the in-house data."},{"cited_title":"Journal of the American College of Cardiology12(2), 315–323 (1988) 4","cited_arxiv_id":null,"evidence_quote":"Grounds the QCA-based computation of %DS from reference diameter and minimum lumen diameter used to edit vessel masks."},{"cited_title":"Circulation 55(2), 329–337 (1977) 4","cited_arxiv_id":null,"evidence_quote":"Defines quantitative coronary angiography measurement principles, including reference diameter and minimal lumen diameter estimates."},{"cited_title":"In: Proceedings of the IEEE/CVF winter conference on applications of computer vision","cited_arxiv_id":null,"evidence_quote":"Demonstrates controllable diffusion augmentation for object detection and serves as the prior work the paper contrasts with its targeted inpainting approach."},{"cited_title":"Journal of the American College of Cardiology79(2), e21–e129 (2022) 2, 4, 6","cited_arxiv_id":null,"evidence_quote":"Provides the clinical severity thresholds (≥50% clinically significant, ≥70% severe) that define the classification labels and synthetic balancing targets."}],"review_version":1}