REVIEW 4 major objections 5 minor 1 cited by
RenderBender: A Survey on Adversarial Attacks Using Differentiable Rendering
T0 review · 4 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read The paper maps every differentiable-rendering adversarial attack into five attacker goals and five scene manipulations, exposing texture attacks as dominant and lighting and sensor attacks as rare.
desk verdict A useful first map of differentiable-rendering attacks, but the survey's own counts don't hold together; fix the corpus and it's worth citing. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The survey's load-bearing object is the crossover of attacker goals and 'scene components' that can be manipulated through the differentiable rendering pipeline. Attacker goals are the five threat-model categories—misclassification, misdetection, reduce confidence, misestimate motion, misestimate depth. Scene manipulations are the five categories of geometry (mesh vertices, point clouds), texture (color, UV maps, reflectance), object pose and translation, illumination (light sources), and sensors (camera and LiDAR parameters). The differentiable renderer is the enabling mechanism: because rendering is differentiable, gradients of the victim DNN's loss flow backward through the renderer to the 3D scene parameters, allowing gradient-based optimization of any manipulated component. The survey's tables cross these two dimensions and use the resulting counts as evidence of which attack surfaces are crowded and which remain open.
What would settle it
A literature search that finds a published differentiable-rendering adversarial attack that does not fit any combination of the five goals and five manipulations—for example, an attack that only alters material reflectance or camera lens distortion—would refute the survey's claim that the framework is exhaustive.
Extended reading notes
Core claim
The central claim is that all existing adversarial attacks using differentiable rendering can be classified by the attacker's goal and by the scene component manipulated, and that this classification is exhaustive and useful. The paper offers a framework with five goals (misclassification, misdetection, reduce confidence, misestimate motion, misestimate depth) and five manipulation categories (geometry, texture, pose, illumination, sensors), then places each of 28 surveyed works in the resulting matrix. The statistical picture that emerges is the paper's main finding: texture manipulation accounts for 15 of the 24 attack papers, geometry for 7, pose for 5, illumination for 2, and sensors for 1; misclassification is the most common goal (18), followed by misdetection (17) and confidence reduction (10), while motion and depth misestimation each have only a single attack. This distribution, the paper argues, reveals that the field has concentrated on visual appearance while neglecting scene-level parameters like lighting and camera settings, and it uses that gap to motivate four future directions: broader target diversity, state-of-the-art models and new modalities, physically realistic phenomena, and better tools and pipelines.
Load-bearing premise
The survey's framework and gap analysis stand on the completeness of its five attacker goals and five scene-manipulation categories and on every surveyed work being an attack correctly assigned to one cell of that matrix.
Editorial extensions
If this is right
- Researchers gain a common coordinate system (goal × manipulated component) for comparing attacks that previously had no shared vocabulary.
- The dominance of texture attacks (15 of 24 attack papers) implies that defensive work against adversarial textures is the most urgent, while illumination and sensor attacks are largely open threat surfaces.
- Because differentiable rendering produces physically plausible scene edits, the surveyed attacks can transfer from digital simulations to physical objects, threatening real-world systems such as autonomous driving.
- The near absence of motion and depth misestimation attacks (one each) suggests those goals are understudied rather than practically safe, and the paper explicitly lists them as future directions.
Reading between the lines
- If the framework is right, the scarcity of illumination and sensor attacks is likely a tooling artifact—current differentiable renderers make whole-scene lighting control difficult—so the grid's lopsidedness may reflect engineering limits rather than the true vulnerability surface.
- A direct testable extension is to build an illumination-only or camera-parameter-only attack with a modern renderer and measure physical transferability; the framework predicts the least-explored cells are the most promising for new attacks.
- As robots and drones adopt NeRF and Gaussian Splatting for scene representation, the same manipulations—for example, rain-like geometry that corrupts optical flow—could be repurposed against embodied perception stacks, not just image classifiers and detectors.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a task-based survey framework for adversarial attacks that use differentiable rendering. It defines five attacker goals (misclassification, misdetection, reduce confidence, misestimate motion, misestimate depth) and five scene-manipulation categories (geometry, texture, pose, illumination, sensors), and it reports a corpus of 28 works classified as 3 surveys, 1 metrics paper, and 24 attacks. The survey also inventories attacked DNN models and attacker access levels, discusses digital and physical attack domains, and proposes future directions centered on target diversity, modern model architectures, real-world phenomena, and tooling.
Significance. A reliable, task-based survey of differentiable-rendering attacks would be a genuinely useful contribution because this literature is fragmented across goals, scene representations, and rendering toolchains. The paper's five-goal/five-manipulation framing is a sensible organizing device, and the inventory of attacked models plus the digital/physical domain discussion add practical value. However, the central claims of the paper are comprehensiveness and the identification of research gaps from quantitative counts, and those claims are not currently reproducible from the tables and text. If the corpus is corrected and the counts recomputed from a consistent per-work category table, the survey could serve as a reference; in its present form the quantitative findings are not trustworthy.
major comments (4)
- [Sec. 2, Table 1] The paper states that the 28 works are categorized as 3 surveys, 1 metrics paper, and 24 attacks, and it derives goal counts from the 24 attacks (18 misclassification, 17 misdetection, 10 reduce confidence, 1 motion, 1 depth). Table 1 as printed is inconsistent with that classification: Wiyatno et al. [2019] is an adversarial-examples survey but is labeled A; Papernot et al. [2016b] is a position paper on the science of security and privacy but is labeled A; and Papernot et al. [2016a] is an early adversarial-examples paper that does not use differentiable rendering but is labeled A. Moreover, Bolya et al. [2020] is an error-analysis toolbox rather than an attack, so the overall corpus is not a clean set of differentiable-rendering attack works. Because these rows enter the attack base or the corpus, the counts and the later claim that texture attacks are the most prevalent (Sec. 4) are not reliable as printed.
- [Table 2] Table 2 lists Hu et al. [2023] and Jiang et al. [2024] as attacked-model columns, but neither work appears in Table 1 or in the stated 28-work corpus. Conversely, eight Table 1 rows (Bolya, Cao, Li et al. 2024b, Machado, Papernot et al. 2016a, Papernot et al. 2016b, Wiyatno, Yuan) have no corresponding column in Table 2. This mismatch means the corpus is not consistently defined across the two central tables, so the comprehensiveness claim in Sec. 1.1 and the research-gap analysis in Sec. 6 must be re-derived after the inventory is reconciled.
- [Sec. 4] The manipulation-category counts in Sec. 4 (texture 15, geometry 7, pose 5, illumination 2, sensors 1) are central to the survey's gap analysis, but the paper does not provide a per-work mapping of the 28 works to these five categories. The colored cells in Table 1 are not enumerated in the text, and after the misclassified rows and missing works are corrected the counts will change. The authors should provide an explicit per-work category table and recompute all counts and gap statements from it.
- [Sec. 1.1] The survey's selection criteria are not stated. The abstract and Sec. 1.1 claim comprehensiveness ('We reviewed 28 works from top venues'), yet no search protocol or inclusion/exclusion criteria are given. Without such a protocol the reader cannot distinguish a deliberate scope restriction from an omission; given the Table 1/Table 2 mismatch and the presence of Hu et al. 2023 and Jiang et al. 2024 in the references, the 'first comprehensive' claim is not defensible as written.
minor comments (5)
- [Abstract] The phrase 'are key ingredient needed to produce' should be 'are key ingredients needed to produce'.
- [Sec. 1.1] The sentence that lists surveys on NeRF and 3D Gaussian Splatting is a run-on and should be split for readability.
- [References] The citation key 'Chen and Wang., 2024' contains a stray period inside the author field and should be reformatted.
- [Sec. 5.2] The quotation formatting around 'sticker-mode' is inconsistent, with an opening quotation mark missing before the word in the text.
- [Table 2] Presenting works as columns and models as rows makes the table hard to scan; transposing it so each row is a work and each column a model would better match the reader's expectations for an attacked-model inventory.
Circularity Check
No circularity found: the survey's taxonomy is definitional organization of existing work, not a derivation whose output is equivalent to its inputs.
full rationale
This is a survey paper; it does not fit parameters, make predictions, or derive new empirical results from assumptions. The five attacker goals (Sec. 2) and five scene-manipulation categories (Secs. 3-4) are organizing categories that the authors apply to 28 existing works in Table 1. That the authors define the taxonomy and then use it to categorize papers is inherent to any classification survey and is not a case where a claimed result reduces to its input by construction. There are no load-bearing self-citations: the cited attack surveys and threat-model sources are external works, and no uniqueness theorem or prior result by the same authors is invoked to force the framework. The internal inconsistencies noted by the reader (Papernot et al. 2016b marked as an attack despite being a position paper; Hu et al. 2023 and Jiang et al. 2024 appearing in Table 2 but not among the 28 Table 1 rows; category counts not reproducible from the printed cells) are reliability and completeness concerns about the survey's inventory, not circularity in the derivation chain. Even if the 'first comprehensive' claim is contestable or the category assignments are wrong, those are correctness issues, not self-referential reasoning. Accordingly, no circular step is identified.
Assumptions & free parameters
assumptions (2)
- domain assumption The 28 works selected are representative of all research on adversarial attacks using differentiable rendering.
- domain assumption The five attacker goals and five scene component categories are exhaustive and non-overlapping.
Cite this review
Pith. "Pith review of RenderBender: A Survey on Adversarial Attacks Using Differentiable Rendering." pith.science (2026). https://pith.science/paper/SSDJRVR2
@misc{pith2026241109749,
author = {Pith},
title = {Pith review of: RenderBender: A Survey on Adversarial Attacks Using Differentiable Rendering},
year = {2026},
howpublished = {\url{https://pith.science/paper/SSDJRVR2}},
note = {Machine review of arXiv:2411.09749}
}
read the original abstract
Differentiable rendering techniques like Gaussian Splatting and Neural Radiance Fields have become powerful tools for generating high-fidelity models of 3D objects and scenes. Their ability to produce both physically plausible and differentiable models of scenes are key ingredient needed to produce physically plausible adversarial attacks on DNNs. However, the adversarial machine learning community has yet to fully explore these capabilities, partly due to differing attack goals (e.g., misclassification, misdetection) and a wide range of possible scene manipulations used to achieve them (e.g., alter texture, mesh). This survey contributes the first framework that unifies diverse goals and tasks, facilitating easy comparison of existing work, identifying research gaps, and highlighting future directions - ranging from expanding attack goals and tasks to account for new modalities, state-of-the-art models, tools, and pipelines, to underscoring the importance of studying real-world threats in complex scenes.
Figures
Forward citations
Cited by 1 Pith paper
-
UNDREAM: Bridging Differentiable Rendering and Photorealistic Simulation for End-to-end Adversarial Attacks
UnDREAM enables optimization of adversarial textures on arbitrary 3D objects inside Unreal Engine by bridging the simulator to the differentiable renderer Mitsuba.
Reference graph
Works this paper leans on
- [2019]
-
[2020]
[Byunet al., 2022 ] J. Byun, S. Cho, M. Kwon, H. Kim, and C. Kim. Improving the Transfer- ability of Targeted Adversarial Examples through Object-Based Diverse Input.arXiv,
work page 2022
-
[2023]
[Huet al., 2023 ] Z. Hu, W. Chu, X. Zhu, H. Zhang, B. Zhang, and X. Hu. Physically Realizable Natural-Looking Clothing Textures Evade Person Detectors via 3D Modeling.CVPR,
work page 2023
-
[2024]
[Donget al., 2022 ] Y . Dong, S. Ruan, H. Su, C. Kang, X. Wei, and J. Zhu. ViewFool: eval- uating the robustness of visual recognition to ad- versarial viewpoints.NeurIPS,
work page 2022
-
[1]
[Abdelfattahet al., 2021 ] M. Abdelfattah, K. Yuan, Z. Wang, and R. Ward. Towards Universal Physi- cal Attacks On Cascaded Camera-Lidar 3d Object Detection Models.ICIP,
work page 2021
-
[6]
[Chen and Wang., 2024] G. Chen and W. Wang. A Survey on 3D Gaussian Splatting.arXiv,
work page 2024
-
[8]
[Gaoet al., 2023 ] K. Gao, Y . Gao, H. He, D. Lu, L. Xu, and J. Li. NeRF: Neural Radiance Field in 3D Vision, A Comprehensive Review.arXiv,
work page 2023
- [10]
Show all 40 references
-
[11]
Jiang, H
[Jianget al., 2024 ] W. Jiang, H. Zhang, X. Wang, Z. Guo, and H. Wang. NeRFail: Neural Radi- ance Fields-Based Multiview Adversarial Attack. AAAI,
2024
-
[12]
[Katoet al., 2020 ] H. Kato, D. Beker, M. Morariu, T. Ando, T. Matsuoka, W.Kehl, and A. Gaidon. Differentiable Rendering: A Survey.arXiv,
2020
-
[13]
Kerbl, G
[Kerblet al., 2023 ] B. Kerbl, G. Kopanas, T. Leimkuehler, and G. Drettakis. 3D Gaussian Splatting for Real-Time Radiance Field Ren- dering.ACM Transactions on Graphics, 42(4),
2023
-
[14]
Leclerc, H
[Leclercet al., 2022 ] G. Leclerc, H. Salman, A. Ilyas, S. Vemprala, L. Engstrom, V . Vineet, K. Xiao, P. Zhang, S. Santurkar, G. Yang, A. Kapoor, and A. Madry. 3DB: A Frame- work for Debugging Computer Vision Models. NeurIPS,
2022
-
[15]
Leheng, Q
[Lehenget al., 2023 ] L. Leheng, Q. Lian, and Y . Chen. Adv3D: Generating 3D Adversarial Ex- amples in Driving Scenarios with NeRF.arXiv,
2023
-
[16]
[Liuet al., 2019 ] H. Liu, M. Tao, C. Li, D. Nowrouzezahrai, and A. Jacobson. Be- yond Pixel Norm-Balls: Parametric Adversaries using an Analytically Differentiable Renderer. ICLR,
2019
-
[17]
Machado, E
[Machadoet al., 2023 ] G. Machado, E. Silva, and R. Goldschmidt. Adversarial Machine Learning in Image Classification: A Survey Toward the Defender’s Perspective.ACM Computing Surveys, 55(1),
2023
-
[18]
Maesumi, M
[Maesumiet al., 2021 ] A. Maesumi, M. Zhu, Y . Wang, T. Chen, Z. Wang, and C. Bajaj. Learning Transferable 3D Adversarial Cloaks for Deep Trained Detectors.arXiv,
2021
-
[19]
Mahendran and A
[Mahendran and Vedaldi, 2015] A. Mahendran and A. Vedaldi. Understanding deep image represen- tations by inverting them.CVPR,
2015
-
[21]
Mildenhall, P
[Mildenhallet al., 2020 ] B. Mildenhall, P. Srini- vasan, M. Tancik, J. Barron, N. Ramamoorthi, and R. Ng. NeRF: Representing Scenes as Neu- ral Radiance Fields for View Synthesis.ECCV,
2020
-
[22]
Miller, Z
[Milleret al., 2020 ] D. Miller, Z. Xiang, and G. Ke- sidis. Adversarial Learning Targeting Deep Neu- ral Network Classification: A Comprehensive Re- view of Defenses Against Attacks.Proceedings of the IEEE, 108(3),
2020
-
[23]
[Mittal, 2024] A. Mittal. Neural Radiance Fields: Past, Present, and Future.arXiv,
2024
-
[25]
Schmalfuss, L
[Schmalfusset al., 2023 ] J. Schmalfuss, L. Mehl, and A. Bruhn. Distracting Downpour: Adversar- ial Weather Attacks for Motion Estimation.ICCV,
2023
-
[26]
Shahreza and S
[Shahreza and Marcel, 2023] H. Shahreza and S. Marcel. Comprehensive Vulnerability Evalua- tion of Face Recognition Systems to Template Inversion Attacks via 3D Face Reconstruction. IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(12),
2023
-
[27]
Sharif, S
[Sharifet al., 2016 ] M. Sharif, S. Bhagavatula, L. Bauer, and M. K. Reiter. Accessorize to a crime: Real and stealthy attacks on state-of-the- art face recognition.CCS,
2016
-
[28]
Suryanto, Y
[Suryantoet al., 2022 ] N. Suryanto, Y . Kim, H. Kang, H. Larasati, Y . Yun, T. Le, H. Yang, S. Oh, and H. Kim. DTA: Physical Camouflage Attacks using Differentiable Transformation Network.CVPR,
2022
-
[29]
Suryanto, Y
[Suryantoet al., 2023 ] N. Suryanto, Y . Kim, H. Larasati, H. Kang, T. Le, Y . Hong, H. Yang, S. Oh, and H. Kim. ACTIVE: Towards Highly Transferable 3D Physical Camouflage for Universal and Robust Vehicle Evasion.ICCV,
2023
-
[30]
Tewari, J
[Tewariet al., 2022 ] A. Tewari, J. Thies, and B. Mildenhall. Advances in Neural Rendering. Computer Graphics F orum, 41(22),
2022
-
[31]
[Tosiet al., 2024 ] F. Tosi, Y . Zhang, Z. Gong, E. Sandström, S. Mattoccia, M.R. Oswald, and M. Poggi. How NeRFs and 3D Gaussian Splatting are Reshaping SLAM: a Survey.arXiv,
2024
-
[32]
[Tuet al., 2021 ] J. Tu, H. Li, X. Yan, M. Ren, Y . Chen, M. Liang, E. Bitar, E. Yumer, and R. Ur- tasun. Exploring Adversarial Robustness of Multi- sensor Perception Systems in Self Driving.CoRL,
2021
-
[33]
[Wanget al., 2022 ] D. Wang, T. Jiang, J. Sun, W. Zhou, X. Zhang, Z. Gong, W. Yao, and X. Chen. FCA: Learning a 3D Full-coverage Vehicle Camouflage for Multi-view Physical Ad- versarial Attack.AAAI,
2022
-
[34]
Wiyatno, A
[Wiyatnoet al., 2019 ] R. Wiyatno, A. Xu, O. Dia, and A. de Berker. Adversarial Examples in Mod- ern Machine Learning: A Review.arXiv,
2019
-
[35]
[Xiaoet al., 2019 ] C. Xiao, D. Yang, B. Li, J. Deng, and M. Liu. MeshAdv: Adversarial Meshes for Visual Recognition.CVPR,
2019
-
[36]
[Xieet al., 2022 ] Y . Xie, T. Takikawa, S. Saito, O. Litany, S. Yan, N. Khan, F. Tombari, J. Tomp- kin, V . Sitzmann, and S. Sridhar. Neural Fields in Visual Computing and Beyond.arXiv,
2022
-
[37]
[Yuanet al., 2019 ] X. Yuan, P. He, Q. Q. Zhu, and X. Li. Adversarial Examples: Attacks and De- fenses for Deep Learning.IEEE Transactions on Neural Networks and Learning Systems, 30(9),
2019
-
[38]
[Zenget al., 2019 ] X. Zeng, C. Liu, Y . Wang, W. Qiu, L. Xie, Y . Tai, C. Tang, and A. Yuille. Ad- versarial Attacks Beyond the Image Space.CVPR,
2019
-
[39]
Zheng, C
[Zhenget al., 2024 ] J. Zheng, C. Lin, J. Sun, Z. Zhao, Q. Li, and C. Shen. Physical 3D Adver- sarial Attacks against Monocular Depth Estima- tion in Autonomous Driving.CVPR,
2024
-
[40]
[Zhouet al., 2024 ] J. Zhou, L. Lyu, D. He, and Y . Li. RAUCA: A Novel Physical Adversarial Attack on Vehicle Detectors via Robust and Ac- curate Camouflage Generation.ICML, 2024
2024
-
[2015]
Meloni, M
[Meloniet al., 2021 ] E. Meloni, M. Tiezzi, L. Pasqualini, M. Gori, and S. Melacci. Messing Up 3D Virtual Environments: Transferable Adversarial 3D Objects.ICMLA,
2021
-
[2016]
Sayles, A
[Sayleset al., 2021 ] A. Sayles, A. Hooda, M. Gupta, R. Chatterjee, and E. Fernandes. Invisible Perturbations: Physical Adversarial Examples Exploiting the Rolling Shutter Effect. CVPR,
2021
-
[2021]
Alcorn, Q
[Alcornet al., 2019 ] M. Alcorn, Q. Li, Z. Gong, C. Wang, L. Mai, W. Ku, and A. Nguyen. Strike (With) a Pose: Neural Networks Are Easily 4https://github.com/mitsuba-renderer/ mitsuba-blender Fooled by Strange Poses of Familiar Objects. CVPR,
2019
-
[2022]
[Caoet al., 2019 ] Y . Cao, C. Xiao, D. Yang, J. Fang, R. Yang, M. Liu, and B. Li. Adversarial Objects Against LiDAR-Based Autonomous Driving Sys- tems.arXiv,
2019
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.