{"id":"e8680f2c-6870-46ae-bea9-332e28dfeb32","arxiv_id":"2606.30115","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":4.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"Adversarial training cuts optimization-based attack success on a U-Net CT model observer to 7% for classification and 13% with localization training, with radiomic texture features linking changes to failures.","lead":"The paper tested a U-Net model for detecting low-contrast objects in CT phantom scans against adversarial attacks and found that dynamic adversarial training reduced attack success rates substantially while preserving performance. A smart generalist might read it to see concrete evidence on making medical AI more reliable against subtle input changes that could affect clinical decisions.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.3","headline":"Phantom dataset and controlled perturbations may not capture clinical CT variability, limiting generalizability of the reported robustness gains.","rationale":"The reader's weakest assumption correctly isolates the key external-validity gap. Because the work is framed as supporting clinical protocol optimization, the phantom-only evaluation is the single condition whose failure would most directly undermine the practical significance of the 7%/13% numbers. All other elements (attack generation, training procedure, ROC confirmation) are internally consistent with the stated experimental scope.","tokens_in":1764,"tokens_out":328,"duration_ms":41556,"concrete_test":"Re-evaluate the adversarially trained U-Net on a second, independent phantom (or on simulated patient CT volumes with added anatomical structures and realistic noise) using the identical optimization-based attack procedure; if success rates rise above 20% while clean-task ROC remains comparable, the headline robustness claim does not generalize beyond the original phantom.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim—that dynamic adversarial training reduces optimization-based attack success to 7% (classification) and 13% (with localization training) without performance loss—rests on results obtained exclusively from a single phantom dataset with low-contrast inserts. Real clinical CT inputs introduce anatomical heterogeneity, beam-hardening, scatter, motion, and scanner-specific noise distributions absent from the phantom. These factors can change both the U-Net feature sensitivities (as hinted by the radiomics texture shifts) and the transferability or effectiveness of the generated perturbations. No evidence is supplied that the 7%/13% figures survive such distribution shift.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript reports an empirical evaluation of adversarial vulnerabilities in a U-Net-based model observer for detecting and localizing low-contrast objects in CT phantom images. Gradient-based attacks achieve up to 75% misclassification while optimization-based attacks reach ~50% success on both tasks; dynamic adversarial training is shown to reduce optimization-based success rates to 7% (classification) and 13% (with localization-specific training) without degrading performance, as confirmed by localization ROC analysis. Radiomic texture features are analyzed to link subtle intensity pattern changes to prediction failures.","tokens_in":1895,"tokens_out":436,"duration_ms":37088,"significance":"If the reported attack-success reductions hold under statistical scrutiny and extend beyond the phantom, the work would supply concrete, actionable evidence that adversarial training can improve robustness of model observers for CT protocol optimization without task-performance penalties. The before/after numerical results and radiomics interpretation are strengths that aid explainability in medical AI.","major_comments":[{"comment":"Abstract: the post-training optimization-attack success rates (7% classification, 13% with localization training) are stated as point estimates with no error bars, trial counts, or statistical tests against the pre-training baselines (~50%), so the magnitude and reliability of the claimed improvement cannot be assessed from the given data.","section":"Abstract"},{"comment":"Abstract and implied Results: all quantitative claims rest on a single phantom dataset with low-contrast inserts. No experiments on clinical CT volumes, multiple phantoms, or acquisition variations (beam-hardening, scatter, motion) are reported, leaving open whether the 7%/13% figures survive the distribution shifts that the skeptic note correctly flags as the weakest assumption.","section":"Abstract"}],"minor_comments":[{"comment":"The abstract refers to 'dynamic adversarial training' and 'localization-specific training' without specifying the training schedule, loss weighting, or hyper-parameters; these details are needed to reproduce the robustness gains.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the thoughtful and constructive review. The comments highlight important aspects of statistical reporting and generalizability. We respond to each major comment below and indicate planned revisions.","responses":[{"response":"We agree that the abstract would benefit from additional statistical context. The reported rates are derived from multiple attack generations on the test set; in the revised manuscript we will specify the number of trials, include error bars (standard deviation across independent runs or bootstrap estimates), and add a brief statement of statistical comparison (e.g., McNemar test) against the pre-training baselines. These details will appear both in the abstract and in an expanded results section.","revision_made":"yes","referee_comment":"[Abstract] Abstract: the post-training optimization-attack success rates (7% classification, 13% with localization training) are stated as point estimates with no error bars, trial counts, or statistical tests against the pre-training baselines (~50%), so the magnitude and reliability of the claimed improvement cannot be assessed from the given data."},{"response":"The work was intentionally performed on a controlled phantom to enable precise, repeatable evaluation of adversarial effects under known imaging conditions, which is standard practice when first characterizing model-observer robustness. We acknowledge that this leaves open questions of robustness under clinical distribution shifts. In the revision we will add an explicit limitations paragraph in the discussion that (i) states the single-phantom scope, (ii) notes the potential impact of untested variations such as beam-hardening or motion, and (iii) outlines planned future validation on clinical or multi-phantom data. The current results therefore constitute a proof-of-concept rather than a claim of broad generalizability.","revision_made":"partial","referee_comment":"[Abstract] Abstract and implied Results: all quantitative claims rest on a single phantom dataset with low-contrast inserts. No experiments on clinical CT volumes, multiple phantoms, or acquisition variations (beam-hardening, scatter, motion) are reported, leaving open whether the 7%/13% figures survive the distribution shifts that the skeptic note correctly flags as the weakest assumption."}],"tokens_in":1422,"tokens_out":494,"duration_ms":42798,"standing_objections":["Whether the reported 7 % / 13 % attack-success reductions persist under realistic clinical distribution shifts cannot be answered without new experiments on clinical volumes or varied acquisition conditions."]},"desk_editor":{"model":"grok-4.3","letter":"This paper shows that dynamic adversarial training brings optimization-based attack success rates down from around 50% to 7% for classification and 13% when localization training is added, all while keeping detection and localization performance steady on their phantom tests.\n\nThey run both gradient-based and optimization attacks, note that localization holds up better than pure classification against small perturbations, and add radiomic texture analysis to show which local intensity features shift in successful attacks. The before-and-after numbers and the unchanged LROC curves give a usable data point for anyone hardening similar model observers.\n\nThe work applies standard white-box methods and adversarial training to a CT protocol optimization task, which is a straightforward extension rather than a new framework. The radiomics step adds a bit of interpretability that is welcome.\n\nThe soft spot is the data. All results come from one phantom with low-contrast inserts. Real clinical CT brings anatomical variation, scatter, motion, and scanner-specific noise that the stress-test note flags correctly; nothing in the abstract tests whether the 7% and 13% figures survive that shift. Error bars and statistical tests are also missing from the summary.\n\nGroups working on robust AI for medical imaging will find the experimental recipe and numbers worth reading. The paper has enough concrete, falsifiable results to deserve peer review, even if reviewers will press on external validity.\n\nI would send it out for review rather than desk reject.","headline":"Adversarial training cuts optimization attack success to 7-13% on this phantom U-Net observer without hurting task performance, but the gains sit on narrow data.","tokens_in":2430,"tokens_out":366,"would_cite":false,"duration_ms":48233,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Dynamic adversarial training reduces optimization-based attack success to 7% for classification and 13% for localization in a U-Net CT model observer without degrading performance.","keywords":["adversarial robustness","U-Net","CT imaging","model observer","adversarial training","low-contrast detection","radiomic analysis","protocol optimization"],"falsifier":"Testing the trained model on real patient CT scans with natural anatomical variations and different perturbation types to check whether the reported attack success rates remain below 13%.","tokens_in":2674,"feed_emoji":"🛡️","tokens_out":606,"duration_ms":34185,"temperature":0.7,"pith_summary":"The paper tests the vulnerability of a U-Net-based model observer to adversarial attacks when detecting and localizing low-contrast objects in CT phantom images for protocol optimization. Gradient-based attacks reach 75% misclassification and optimization-based attacks reach about 50% success on both tasks. Dynamic adversarial training lowers those rates sharply while radiomic texture analysis shows that successful attacks alter specific local intensity features. A reader cares because this points to a practical way to make AI tools in medical imaging more reliable against small input changes. The approach keeps localization receiver operating characteristic performance intact.","feed_headline":"Adversarial training cuts CT model attack success to 7%","feed_subtitle":"Dynamic training on U-Net observer protects low-contrast object detection and localization without performance loss","key_machinery":"Dynamic adversarial training applied to the U-Net-based model observer, which iteratively generates perturbations during training to build resistance to both classification and localization failures.","core_discovery":"Adversarial attacks generated with gradient-based and optimization-based white-box methods expose vulnerabilities in the U-Net model observer, but dynamic adversarial training reduces the success rate of optimization-based attacks to 7% for classification and 13% when including localization-specific training, without compromising task performances as confirmed by localization receiver operating characteristic analysis.","pith_inferences":["Similar dynamic training could be tested on other segmentation or detection networks used in medical imaging beyond CT.","The observed sensitivity to local texture changes suggests adding explicit texture-regularization terms during training as a next step.","Extending the evaluation to multi-center clinical datasets would reveal whether phantom-based robustness transfers to varied scanner protocols."],"forward_implications":["The model observer achieves substantially lower success rates for both gradient-based and optimization-based attacks after training.","Detection and localization performance stay intact as shown by unchanged receiver operating characteristic curves.","Radiomic texture features provide an interpretable link between image alterations and prediction failures.","The method supports development of more reliable AI for CT protocol optimization tasks."],"fun_headline_variants":["Adversarial training reduces U-Net CT attack success to 7%","U-Net CT observer more resistant after dynamic training","White-box attacks on CT model mitigated by adversarial training","Adversarial training protects U-Net detection in low-contrast CT"],"cache_read_input_tokens":64,"weakest_assumption_plain":"The phantom dataset and generated adversarial perturbations represent the range of input variations and threats found in real clinical CT imaging.","fun_headline_variants_meta":{"raw":{"variants":["Adversarial training reduces U-Net CT attack success to 7%","U-Net CT observer more resistant after dynamic training","White-box attacks on CT model mitigated by adversarial training","Adversarial training protects U-Net detection in low-contrast CT"]},"model":"grok-4.3","cost_usd":0.009209,"raw_usage":{"total_tokens":4140,"prompt_tokens":698,"num_sources_used":0,"completion_tokens":65,"cost_in_usd_ticks":92087000,"prompt_tokens_details":{"text_tokens":698,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":3377,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":698,"tokens_out":65,"duration_ms":48127,"temperature":1.0,"reasoning_tokens":3377,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-30T04:04:53.610381+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Testing the trained model on real patient CT scans with natural anatomical variations and different perturbation types to check whether the reported attack success rates remain below 13%.","supporting_citations":[],"review_version":1}