REVIEW 2 major objections 1 minor 53 references
Adversarial robustness of a U-Net-based model observer for CT protocol optimization
T0 review · 2 major / 1 minor · reviewed 2026-06-30 · grok-4.3
Pith's one-line read Dynamic adversarial training reduces optimization-based attack success to 7% for classification and 13% for localization in a U-Net CT model observer without degrading performance.
desk verdict Adversarial training cuts optimization attack success to 7-13% on this phantom U-Net observer without hurting task performance, but the gains sit on narrow data. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Dynamic adversarial training applied to the U-Net-based model observer, which iteratively generates perturbations during training to build resistance to both classification and localization failures.
What would settle it
Testing the trained model on real patient CT scans with natural anatomical variations and different perturbation types to check whether the reported attack success rates remain below 13%.
Extended reading notes
Core claim
Adversarial attacks generated with gradient-based and optimization-based white-box methods expose vulnerabilities in the U-Net model observer, but dynamic adversarial training reduces the success rate of optimization-based attacks to 7% for classification and 13% when including localization-specific training, without compromising task performances as confirmed by localization receiver operating characteristic analysis.
Load-bearing premise
The phantom dataset and generated adversarial perturbations represent the range of input variations and threats found in real clinical CT imaging.
Editorial extensions
If this is right
- The model observer achieves substantially lower success rates for both gradient-based and optimization-based attacks after training.
- Detection and localization performance stay intact as shown by unchanged receiver operating characteristic curves.
- Radiomic texture features provide an interpretable link between image alterations and prediction failures.
- The method supports development of more reliable AI for CT protocol optimization tasks.
Reading between the lines
- Similar dynamic training could be tested on other segmentation or detection networks used in medical imaging beyond CT.
- The observed sensitivity to local texture changes suggests adding explicit texture-regularization terms during training as a next step.
- Extending the evaluation to multi-center clinical datasets would reveal whether phantom-based robustness transfers to varied scanner protocols.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript reports an empirical evaluation of adversarial vulnerabilities in a U-Net-based model observer for detecting and localizing low-contrast objects in CT phantom images. Gradient-based attacks achieve up to 75% misclassification while optimization-based attacks reach ~50% success on both tasks; dynamic adversarial training is shown to reduce optimization-based success rates to 7% (classification) and 13% (with localization-specific training) without degrading performance, as confirmed by localization ROC analysis. Radiomic texture features are analyzed to link subtle intensity pattern changes to prediction failures.
Significance. If the reported attack-success reductions hold under statistical scrutiny and extend beyond the phantom, the work would supply concrete, actionable evidence that adversarial training can improve robustness of model observers for CT protocol optimization without task-performance penalties. The before/after numerical results and radiomics interpretation are strengths that aid explainability in medical AI.
major comments (2)
- [Abstract] Abstract: the post-training optimization-attack success rates (7% classification, 13% with localization training) are stated as point estimates with no error bars, trial counts, or statistical tests against the pre-training baselines (~50%), so the magnitude and reliability of the claimed improvement cannot be assessed from the given data.
- [Abstract] Abstract and implied Results: all quantitative claims rest on a single phantom dataset with low-contrast inserts. No experiments on clinical CT volumes, multiple phantoms, or acquisition variations (beam-hardening, scatter, motion) are reported, leaving open whether the 7%/13% figures survive the distribution shifts that the skeptic note correctly flags as the weakest assumption.
minor comments (1)
- [Abstract] The abstract refers to 'dynamic adversarial training' and 'localization-specific training' without specifying the training schedule, loss weighting, or hyper-parameters; these details are needed to reproduce the robustness gains.
Simulated Author's Rebuttal
We thank the referee for the thoughtful and constructive review. The comments highlight important aspects of statistical reporting and generalizability. We respond to each major comment below and indicate planned revisions.
read point-by-point responses
-
Referee: [Abstract] Abstract: the post-training optimization-attack success rates (7% classification, 13% with localization training) are stated as point estimates with no error bars, trial counts, or statistical tests against the pre-training baselines (~50%), so the magnitude and reliability of the claimed improvement cannot be assessed from the given data.
Authors: We agree that the abstract would benefit from additional statistical context. The reported rates are derived from multiple attack generations on the test set; in the revised manuscript we will specify the number of trials, include error bars (standard deviation across independent runs or bootstrap estimates), and add a brief statement of statistical comparison (e.g., McNemar test) against the pre-training baselines. These details will appear both in the abstract and in an expanded results section. revision: yes
-
Referee: [Abstract] Abstract and implied Results: all quantitative claims rest on a single phantom dataset with low-contrast inserts. No experiments on clinical CT volumes, multiple phantoms, or acquisition variations (beam-hardening, scatter, motion) are reported, leaving open whether the 7%/13% figures survive the distribution shifts that the skeptic note correctly flags as the weakest assumption.
Authors: The work was intentionally performed on a controlled phantom to enable precise, repeatable evaluation of adversarial effects under known imaging conditions, which is standard practice when first characterizing model-observer robustness. We acknowledge that this leaves open questions of robustness under clinical distribution shifts. In the revision we will add an explicit limitations paragraph in the discussion that (i) states the single-phantom scope, (ii) notes the potential impact of untested variations such as beam-hardening or motion, and (iii) outlines planned future validation on clinical or multi-phantom data. The current results therefore constitute a proof-of-concept rather than a claim of broad generalizability. revision: partial
- Whether the reported 7 % / 13 % attack-success reductions persist under realistic clinical distribution shifts cannot be answered without new experiments on clinical volumes or varied acquisition conditions.
Circularity Check
Purely empirical study; no derivation chain or self-referential reductions present.
full rationale
The paper reports measured attack success rates and robustness improvements from adversarial training on a held-out phantom dataset. No equations, first-principles derivations, fitted parameters renamed as predictions, or self-citation chains appear in the abstract or described content. All claims rest on direct experimental outcomes rather than any reduction to inputs by construction. This matches the default expectation of no significant circularity for non-derivational empirical work.
Assumptions & free parameters
Cite this review
Pith. "Pith review of Adversarial robustness of a U-Net-based model observer for CT protocol optimization." pith.science (2026). https://pith.science/paper/E4LAMBR4
@misc{pith2026260630115,
author = {Pith},
title = {Pith review of: Adversarial robustness of a U-Net-based model observer for CT protocol optimization},
year = {2026},
howpublished = {\url{https://pith.science/paper/E4LAMBR4}},
note = {Machine review of arXiv:2606.30115}
}
read the original abstract
Artificial intelligence is increasingly used in medical imaging, yet its robustness to input perturbations remains a critical concern for a wide clinical adoption. To this end, we used adversarial examples to systematically probe vulnerabilities of a U-Net-based model observer for computed tomography protocol optimization, performing detection and localization of low-contrast objects in a phantom dataset. Adversarial attacks were generated using both gradient-based and optimization-based white-box methods. Fast gradient perturbations produced high misclassification rates, reaching up to 75% at intermediate perturbation levels while remaining visually imperceptible. Localization was more robust, with success rates of about 25% for small perturbations and 42% at moderate levels. In contrast, optimization-based attack achieved success rates close to 50% for both tasks. To mitigate these vulnerabilities, dynamic adversarial training was implemented. This reduced the success rate of optimization-based attacks to 7% for classification and 13% when including localization-specific training, demonstrating a substantial robustness improvement without compromising task performances, confirmed by localization receiver operating characteristic analysis. To further interpret model behavior, radiomic texture analysis was performed on original and adversarial images. While most global image statistics remain stable, specific texture-related features exhibit consistent changes in successful attacks, highlighting the model's sensitivity to subtle local intensity patterns. Overall, adversarial training improves robustness without degrading performance, while radiomic analysis reveals interpretable links between texture alterations and prediction failures, supporting more reliable and explainable AI systems for medical imaging.
Figures
Figures from the paper (8 more)
Reference graph
Works this paper leans on
-
[1]
Hugo J. W. L. Aerts, Emmanuel Rios Velazquez, Ralph T. H. Leijenaar, Chintan Parmar, Patrick Grossmann, Sara Carvalho, Johan Bussink, Ren´ e Monshouwer, Benjamin Haibe-Kains, Derek Ri- 14 etveld, et al. Decoding tumour phenotype by noninvasive imaging using a quantitative radiomics approach.Nature Communications, 5:4006, 2014
2014
-
[2]
Threat of adversarial attacks on deep learning in computer vision: a survey.IEEE Access, 6:14410–14430, 2018
Naveed Akhtar and Ajmal Mian. Threat of adversarial attacks on deep learning in computer vision: a survey.IEEE Access, 6:14410–14430, 2018
2018
-
[3]
Adversarial attacks and defenses in explainable artificial intelligence: a survey.Information Fusion, 107:102303, 2024
Hubert Baniecki and Przemys law Biecek. Adversarial attacks and defenses in explainable artificial intelligence: a survey.Information Fusion, 107:102303, 2024
2024
-
[4]
Barrett and Kyle J
Harrison H. Barrett and Kyle J. Myers.Foundations of Image Science. Wiley, 2004
2004
-
[5]
Barrett, Jie Yao, Jannick P
Harrison H. Barrett, Jie Yao, Jannick P. Rolland, and Kyle J. Myers. Model observers for assessment of image quality.Proceedings of the National Academy of Sciences, 90(21):9758–9765, 1993
1993
-
[6]
Evasion attacks against machine learning at test time
Battista Biggio, Igino Corona, Davide Maiorca, Blaine Nelson, Nedim Srndic, Pavel Laskov, Giorgio Giacinto, and Fabio Roli. Evasion attacks against machine learning at test time. InECML PKDD, pages 387–402, 2013
2013
-
[7]
Machine learning robustness: A primer
Houssem Ben Braiek and Foutse Khomh. Machine learning robustness: A primer. InTrustworthy AI in Medical Imaging, pages 37–71. Elsevier, 2025
2025
-
[8]
C. Buckner. Understanding adversarial examples requires a theory of artefacts for deep learning. Nature Machine Intelligence, 2:731–736, 2020
2020
Show all 53 references
-
[9]
Towards evaluating the robustness of neural networks
Nicholas Carlini and David Wagner. Towards evaluating the robustness of neural networks. In2017 IEEE Symposium on Security and Privacy, pages 39–57, 2017
2017
-
[10]
A survey on adversarial attacks and defences.CAAI Transactions on Intelligence Technology, 6(1):25–45, 2021
Anirban Chakraborty, Manaar Alam, Vishal Dey, Anupam Chattopadhyay, and Debdeep Mukhopadhyay. A survey on adversarial attacks and defences.CAAI Transactions on Intelligence Technology, 6(1):25–45, 2021
2021
-
[11]
Council directive 2013/59/euratom, 2013
Council of the European Union. Council directive 2013/59/euratom, 2013
2013
-
[12]
DeGrave, Joseph D
Alex J. DeGrave, Joseph D. Janizek, and Su-In Lee. AI for radiographic covid-19 detection selects shortcuts over signal.Nature Machine Intelligence, 3:610–619, 2021
2021
-
[13]
Survey on adversarial attack and defense for medical image analysis: Methods and challenges.ACM Comput
Junhao Dong, Junxi Chen, Xiaohua Xie, Jianhuang Lai, and Hao Chen. Survey on adversarial attack and defense for medical image analysis: Methods and challenges.ACM Comput. Surv., 57(3):1–38, 2024
2024
-
[14]
Addressing signal alterations induced in CT images by deep learning processing: a preliminary phantom study.Physica Medica, 83:88–100, 2021
Sandra Doria, Federico Valeri, Lorenzo Lasagni, Valentina Sanguineti, Ruggero Ragonesi, Muham- mad Usman Akbar, Alessio Gnerucci, Alessio Del Bue, Alessandro Marconi, Guido Risaliti, et al. Addressing signal alterations induced in CT images by deep learning processing: a preli...
2021
-
[15]
Finlayson, John D
Samuel G. Finlayson, John D. Bowers, Joichi Ito, Jonathan L. Zittrain, Andrew L. Beam, and Isaac S. Kohane. Adversarial attacks on medical machine learning.Science, 363(6433):1287–1289, 2019
2019
-
[16]
Freiesleben and T
T. Freiesleben and T. Grote. Beyond generalization: a theory of robustness in machine learning. Synthese, 202, 2023
2023
-
[17]
The intriguing relation between counterfactual explanations and adversarial examples.Minds and Machines, 32(1):77–109, 2022
Timo Freiesleben. The intriguing relation between counterfactual explanations and adversarial examples.Minds and Machines, 32(1):77–109, 2022
2022
-
[18]
Wichmann
Robert Geirhos, J¨ orn-Henrik Jacobsen, Claudio Michaelis, Richard Zemel, Wieland Brendel, Matthias Bethge, and Felix A. Wichmann. Shortcut learning in deep neural networks.Nature Machine Intelligence, 2(11):665–673, 2020
2020
-
[19]
Gillies, Paul E
Robert J. Gillies, Paul E. Kinahan, and Hedvig Hricak. Radiomics: images are more than pictures, they are data.Radiology, 278(2):563–577, 2016
2016
-
[20]
Goodfellow, Jonathon Shlens, and Christian Szegedy
Ian J. Goodfellow, Jonathon Shlens, and Christian Szegedy. Explaining and harnessing adversarial examples.arXiv preprint arXiv:1412.6572, 2014. 15
2014 arXiv
-
[21]
Haralick, K
Robert M. Haralick, K. Shanmugam, and Its’hak Dinstein. Textural features for image classification. IEEE Transactions on Systems, Man, and Cybernetics, SMC-3(6):610–621, 1973
1973
-
[22]
Model observers in medical imaging research.Theranostics, 3:774–786, 2013
Xin He and Subok Park. Model observers in medical imaging research.Theranostics, 3:774–786, 2013
2013
-
[23]
Hirano, A
H. Hirano, A. Minagi, and K. Takemoto. Universal adversarial attacks on deep neural networks for medical image classification.BMC Medical Imaging, 21(9), 2021
2021
-
[24]
Adversarial examples are not bugs, they are features
Andrew Ilyas, Shibani Santurkar, Dimitris Tsipras, Logan Engstrom, Brandon Tran, and Aleksander Madry. Adversarial examples are not bugs, they are features. InAdvances in Neural Information Processing Systems, 2019
2019
-
[25]
Recommendations of the ICRP
International Commission on Radiological Protection. Recommendations of the ICRP. ICRP pub- lication 26.Ann. ICRP, 1(3), 1977
1977
-
[26]
Artificial intelligence (AI) – assessment of the robustness of neural networks – part 1: Overview
ISO/IEC. Artificial intelligence (AI) – assessment of the robustness of neural networks – part 1: Overview. Technical Report TR 24029-1, ISO/IEC, 2021
2021
-
[27]
Adversarial examples in the physical world
Alexey Kurakin, Ian Goodfellow, and Samy Bengio. Adversarial examples in the physical world. arXiv preprint arXiv:1607.02533, 2016
2016 arXiv
-
[28]
Towards deep learning models resistant to adversarial attacks
Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu. Towards deep learning models resistant to adversarial attacks. InInternational Conference on Learning Representations (ICLR), 2018
2018
-
[29]
Mayerhoefer, Andrzej Materka, Georg Langs, Ida H¨ aggstr¨ om, Piotr Szczypi´ nski, Peter Gibbs, and Gary Cook
Marius E. Mayerhoefer, Andrzej Materka, Georg Langs, Ida H¨ aggstr¨ om, Piotr Szczypi´ nski, Peter Gibbs, and Gary Cook. Introduction to radiomics.Journal of Nuclear Medicine, 61(4):488–495, 2020
2020
-
[30]
Leanpub, 2025
Christoph Molnar.Interpretable Machine Learning. Leanpub, 2025
2025
-
[31]
Universal adversarial perturbations
Seyed-Mohsen Moosavi-Dezfooli, Alhussein Fawzi, Omar Fawzi, and Pascal Frossard. Universal adversarial perturbations. InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 86–94, 2017
2017
-
[32]
Stacked hourglass networks for human pose estimation
Alejandro Newell, Kaiyu Yang, and Jia Deng. Stacked hourglass networks for human pose estimation. InEuropean Conference on Computer Vision (ECCV), pages 483–499, 2016
2016
-
[33]
Adversarial Robustness Toolbox v1.0.0, 2018
Maria-Irina Nicolae, Mathieu Sinn, Minh Ngoc Tran, Beat Buesser, Ambrish Rawat, Martin Wis- tuba, Valentina Zantedeschi, Nathalie Baracaldo, Bryant Chen, Heiko Ludwig, Ian Molloy, and Ben Edwards. Adversarial Robustness Toolbox v1.0.0, 2018
2018
-
[34]
Generalizability vs robustness: investigating medical imaging networks using adversarial examples
Magdalini Paschali, Sailesh Conjeti, Fernando Navarro, and Nassir Navab. Generalizability vs robustness: investigating medical imaging networks using adversarial examples. InMedical Image Computing and Computer-Assisted Intervention (MICCAI), pages 493–501, 2018
2018
-
[35]
Integrating spatial configuration into heatmap regression cnns for landmark localization.Medical Image Analysis, 54:207–219, 2019
Christian Payer, Darko ˇStern, Horst Bischof, and Martin Urschler. Integrating spatial configuration into heatmap regression cnns for landmark localization.Medical Image Analysis, 54:207–219, 2019
2019
-
[36]
On the relationship between generalization and robustness to adversarial examples.Symmetry, 13(5):817, 2021
Anibal Pedraza et al. On the relationship between generalization and robustness to adversarial examples.Symmetry, 13(5):817, 2021
2021
-
[37]
Adversarial machine learning: a review of methods, tools, and critical industry sectors.Artificial Intelligence Review, 2025
Sotiris Pelekis, Thanos Koutroubas, Afroditi Blika, Anastasis Berdelis, Evangelos Karakolis, Chris- tos Ntanos, Evangelos Spiliotis, and Dimitris Askounis. Adversarial machine learning: a review of methods, tools, and critical industry sectors.Artificial Intelligence Review, 2025
2025
-
[38]
U-Net: Convolutional networks for biomedical image segmentation
Olaf Ronneberger, Philipp Fischer, and Thomas Brox. U-Net: Convolutional networks for biomedical image segmentation. InMedical Image Computing and Computer-Assisted Intervention (MICCAI), Lecture Notes in Computer Science, pages 234–241. Springer, 2015
2015
-
[39]
Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Batra
Ramprasaath R. Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Batra. Grad-cam: visual explanations from deep networks via gradient-based local- ization. InProceedings of the IEEE International Conference on Computer Vision (ICCV), pages...
2017
-
[40]
Adversarial examples—a complete characterisation of the phenomenon.arXiv preprint arXiv:1810.01185, 2018
Alexandru Constantin Serban, Erik Poll, and Joost Visser. Adversarial examples—a complete characterisation of the phenomenon.arXiv preprint arXiv:1810.01185, 2018
2018 arXiv
-
[41]
Elkin, and Vijay Devabhaktuni
Nahian Siddique, Sidike Paheding, Colin P. Elkin, and Vijay Devabhaktuni. U-Net and its variants for medical image segmentation: A review of theory and applications.IEEE Access, 9:82031–82057, 2021
2021
-
[42]
Swensson
Richard G. Swensson. Unified measurement of observer performance in detecting and localizing target objects on images.Medical Physics, 23(10):1709–1725, 1996
1996
-
[43]
Intriguing properties of neural networks.arXiv preprint arXiv:1312.6199, 2014
Christian Szegedy et al. Intriguing properties of neural networks.arXiv preprint arXiv:1312.6199, 2014
2014 arXiv
-
[44]
Thibault, B
G. Thibault, B. Fertil, C. Navarro, S. Pereira, P. Cau, N. Levy, J. Sequeira, and J.-L. Mari. Texture indexes and gray level size zone matrix. InProceedings of PRIP 2009, pages 140–145, 2009
2009
-
[45]
Robustness may be at odds with accuracy
Dimitris Tsipras, Shibani Santurkar, Logan Engstrom, Alexander Turner, and Aleksander Madry. Robustness may be at odds with accuracy. InInternational Conference on Learning Representations (ICLR), 2019
2019
-
[46]
Federico Valeri, Maurizio Bartolucci, Elena Cantoni, Roberto Carpi, Evaristo Cisbani, Ilaria Cup- paro, Sandra Doria, Cesare Gori, Mauro Grigioni, Lorenzo Lasagni, et al. U-Net and MobileNet CNN-based model observers for CT protocol optimization: comparative performance evalua...
2023
-
[47]
Computational radiomics system to decode the radiographic phenotype.Cancer Research, 77(21):e104–e107, 2017
Joost JM Van Griethuysen et al. Computational radiomics system to decode the radiographic phenotype.Cancer Research, 77(21):e104–e107, 2017
2017
-
[48]
van Timmeren, Davide Cester, Stephanie Tanadini-Lang, Hatem Alkadhi, and Bettina Baessler
Janita E. van Timmeren, Davide Cester, Stephanie Tanadini-Lang, Hatem Alkadhi, and Bettina Baessler. Radiomics in medical imaging: How-to guide and critical reflection.Insights into Imaging, 11(1):91, 2020
2020
-
[49]
Counterfactual explanations without opening the black box: automated decisions and the gdpr.Harvard Journal of Law & Technology, 31(2):841– 887, 2018
Sandra Wachter, Brent Mittelstadt, and Chris Russell. Counterfactual explanations without opening the black box: automated decisions and the gdpr.Harvard Journal of Law & Technology, 31(2):841– 887, 2018
2018
-
[50]
Welch, Chris McIntosh, Benjamin Haibe-Kains, Michael F
Mattea L. Welch, Chris McIntosh, Benjamin Haibe-Kains, Michael F. Milosevic, Leonard Wee, Andre Dekker, Shao Hui Huang, Thomas G. Purdie, Brian O’Sullivan, Hugo J. W. L. Aerts, and David A. Jaffray. Vulnerabilities of radiomic signature development: the need for safeguards.Ra-...
2019
-
[51]
Han Xu, Yao Ma, Hao-Chen Liu, Debayan Deb, Hui Liu, Ji-Liang Tang, and Anil K. Jain. Adversar- ial attacks and defenses in images, graphs and text: a review.International Journal of Automation and Computing, 17(2):151–178, 2020
2020
-
[52]
Xing, Laurent El Ghaoui, and Michael I
Hongyang Zhang, Yaodong Yu, Jiantao Jiao, Eric P. Xing, Laurent El Ghaoui, and Michael I. Jordan. Theoretically principled trade-off between robustness and accuracy. InInternational Conference on Machine Learning (ICML), pages 7472–7482, 2019
2019
-
[53]
Abdalah, Hugo J
Alex Zwanenburg, Martin Vallieres, Mahmoud A. Abdalah, Hugo J. W. L. Aerts, Vincent Andrea- rczyk, Aditya Apte, Saeed Ashrafinia, Spyridon Baez, Roelof Bakr, et al. The image biomarker standardization initiative: standardized quantitative radiomics for high-throughput image-ba...
2020
Reviewed June 30, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.