{"id":"853b226f-fbd6-4476-9e44-b76d73a71220","arxiv_id":"2505.16402","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"AdvReal generates clothing textures via joint 2D-3D adversarial training with non-rigid cloth and lighting simulation, reporting higher attack success against pedestrian detectors than prior patch methods.","lead":"This paper introduces AdvReal, a method that prints adversarial patterns onto clothing to make AI pedestrian detectors fail to see the wearer. It combines 2D and 3D training with simulated fabric folds and lighting changes, and reports higher physical attack success than prior patch methods.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The abstract's central physical claim (70.13% ASR on YOLOv12 in physical scenarios) is actually the digital closed-box result from Table 3; the physical experiments in §4.7 test only YOLOv5 and lack a clean-clothing control and error bars.","rationale":"The reader's weakest_assumption focused on the sim-to-real fidelity of the 3D rendering pipeline. That is a genuine concern, but the more immediate and checkable problem is that the signature physical claim is not a physical result: the 70.13% YOLOv12 ASR and the baseline numbers it is compared against come from the digital closed-box evaluation in Table 3, while the physical experiments in Section 4.7 use YOLOv5 as the victim model. This is a reporting/verification failure rather than merely an assumption about render realism. I also note the physical protocol lacks a no-patch control and error bars, so even the YOLOv5 physical ASRs cannot be cleanly attributed to the adversarial texture. These issues do not invalidate the digital contributions or the overall engineering effort, and the reader's CONDITIONAL verdict already captures the need for correction and additional evidence. I therefore recommend no change to the verdict, while emphasizing that the abstract's physical wording should be revised unless the requested YOLOv12 physical evaluation with controls can be provided.","tokens_in":23205,"tokens_out":5597,"duration_ms":47453,"concrete_test":"From the released repository, reconstruct the physical evaluation: run the recorded videos through YOLOv12 (the claimed victim) and YOLOv5, with per-frame annotations for distance, yaw angle, lighting, and subject identity; add a clean-clothing/no-patch control at 2m, 3m, and 4m; compute ASR with 95% confidence intervals for each condition. If the 70.13% YOLOv12 figure is only the digital Table 3 entry, and the claimed >90% ASR for frontal and oblique views at 4m cannot be reproduced with controls and intervals, the abstract's physical claim must be corrected to the digital closed-box setting or the physical experiments must be rerun on YOLOv12.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative claim in the abstract is not backed by the reported experimental record. The 70.13% ASR on YOLOv12, and the T-SEA 21.65% / AdvTexture 19.70% comparisons, are digital closed-box numbers from Table 3, not physical results. Section 4.7.2 explicitly states that the physical distance experiments use YOLOv5 as the victim model, and Tables 10-11 contain no YOLOv12 row. The separate claim of '>90% under frontal and oblique views at 4m' also does not appear as disaggregated angle-by-distance data anywhere in Section 4.7. Additionally, the physical tables include no clean-clothing or non-adversarial control condition, so ASR at 4m may include natural detection failures at small scale. The per-cell sample sizes appear to be about 37 frames (e.g., 35.14% = 13/37, 81.08% = 30/37), yielding confidence intervals wide enough to affect the reported ordering, yet no intervals are given. Thus the central claim that AdvReal robustly evades YOLOv12 in physical scenarios currently rests on an unstated extrapolation from digital evaluation plus an underspecified physical protocol.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes AdvReal, a framework for generating physical adversarial patches against pedestrian detectors. The method jointly optimizes adversarial textures in 2D image space and 3D mesh space, incorporating physics-inspired non-rigid cloth deformation (Eqs. 2-8), a time-space/relighting mapping module (Section 3.4, Algorithm 1), and ShakeDrop-style stochastic gradient regularization. The authors evaluate on the INRIA and nuScenes datasets, compare against AdvPatch, AdvTshirt, NatPatch, AdvTexture, AdvCaT, and T-SEA, and report both digital and physical experiments. The abstract claims a 70.13% ASR on YOLOv12 in physical scenarios and >90% ASR at 4 meters under frontal and oblique views.","tokens_in":23389,"tokens_out":7724,"duration_ms":58078,"significance":"If the physical claims were properly supported, the paper would make a useful contribution to adversarial robustness evaluation for autonomous driving perception: the joint 2D-3D optimization with explicit cloth deformation and relighting is a plausible way to improve physical transferability, and the digital experiments show large margins over baselines (e.g., 70.13% vs. 21.65% for T-SEA on YOLOv12 in Table 3). The authors also provide code and a demo video, which aids reproducibility. However, the physical evidence for the central claim is currently weak: the physical experiments use only YOLOv5 as the victim model, lack a clean-clothing control, and report no confidence intervals or per-cell trial counts. The current manuscript does not substantiate the abstract's physical YOLOv12 claim or the >90% at-4m claim.","major_comments":[{"comment":"The abstract states that AdvReal achieves 'an average attack success rate (ASR) of 70.13% on YOLOv12 in physical scenarios' and exceeds 90% under frontal and oblique views at 4 meters. However, the 70.13% ASR appears in Table 3, which is the digital closed-box experiment, not a physical experiment. The physical experiments in Tables 10-11 do not include any YOLOv12 row, and Section 4.7.2 explicitly states that the victim model was YOLOv5. No angle-by-distance disaggregated physical data supporting the >90% at-4m claim is presented anywhere in Section 4.7. This mismatch between the abstract and the reported experimental record must be resolved, either by removing the unsupported physical YOLOv12 claim or by providing the corresponding physical experiments.","section":"Abstract; Section 4.2.1 (Table 3); Section 4.7.2 (Tables 10-11)"},{"comment":"The physical distance experiments lack a clean-clothing or non-adversarial control condition, so the reported ASR values may include natural detection failures at small scale (especially at 4 m), and the contribution of the adversarial patch to the miss rate is not identified. In addition, the text says 'Each patch was taken 111 times at different distances,' but the percentages in Tables 10-11 imply denominators of only about 37 per distance cell (e.g., 81.08% = 30/37, 35.14% = 13/37). No confidence intervals or per-cell sample sizes are reported. Without such statistical reporting, the observed ordering of methods in the physical tables is not supported.","section":"Section 4.7.2 (Tables 10-11)"},{"comment":"Section 4.1.2 states that 'we conduct closed-box attack tests on SOTA detectors (YOLO-v8, v11, v12) in the physical world for the first time,' but the physical experiments in Section 4.7.2 only use YOLOv5 as the victim model. Tables 10-11 list patches trained on different detectors (YOLOv2, YOLOv3, YOLOv5, F-RCNN, D-DETR) but evaluate all of them with YOLOv5. The claim about physical closed-box tests on YOLOv8/v11/v12 is therefore not supported. This should be corrected to avoid overstating the physical evaluation.","section":"Section 4.1.2; Section 4.7.2"},{"comment":"In Table 7, the F1-score reported for AdvReal on YOLOv8 is 70.26%, but the precision (56.55%) and recall (32.68%) in the same row imply F1 = 2·56.55·32.68/(56.55+32.68) ≈ 41.4%, not 70.26%. This numerical inconsistency affects the transferability comparison and should be corrected. The corresponding text in Section 4.5.3 appears to rely on these numbers.","section":"Table 7"}],"minor_comments":[{"comment":"Equation (6) appears garbled: '|C|=max{Nmax, j ρ|S|/ko}' contains undefined symbols (j, k, o), and the text immediately before it refers to N_min while the equation uses N_max. Please clarify the intended formula and the roles of all parameters.","section":"Section 3.3, Eq. (6)"},{"comment":"In Algorithm 1, the update line 'ε←ε−η∇εLtot, ε=[α,β,θ]' is followed by clipping only α and β, and the condition 'α<[α_l,α_h]' is not a valid comparison. The algorithm would be clearer with explicit bounds checks for each parameter, including θ.","section":"Algorithm 1"},{"comment":"Table 5 reports the ablation results but does not state which detector, dataset, or evaluation protocol (digital or physical, glass-box or closed-box) is used. The column header 'AC↑' is also inconsistent with the text, which describes lower AC as better.","section":"Table 5, caption and text"},{"comment":"The term 'closed-box' is used to describe the attack in Section 3.5.3, where the detector's gradients are used for optimization. This is white-box access by standard terminology; the later use of 'closed-box' to mean a held-out detector not used in training (Section 4.1.5) should be distinguished, or a different term such as 'black-box transfer' should be used.","section":"Section 3.2.3, Section 3.5.3"},{"comment":"The sentence 'For the adversarial patches of this paper trained with different target detectors. DDETR is used as a target detector of the transfromer architecture.' is grammatically incomplete and contains typos ('transfromer'). It should be rewritten for clarity.","section":"Section 4.7.2"},{"comment":"The caption of Figure 9 does not state whether the angle-dependent results are from physical or digital experiments, or which victim detector was used. This should be clarified to avoid ambiguity.","section":"Figure 9 caption"}],"recommendation":"major_revision","confidential_remarks":"The paper's digital results are strong and the methodological idea (joint 2D-3D optimization with non-rigid deformation and relighting) is timely. However, the abstract's physical YOLOv12 claim is contradicted by the experimental record, and the physical evaluation as reported is statistically under-powered and lacks controls. I believe these issues are correctable—by revising the abstract, adding proper physical experiments with the claimed detectors, and including controls and error bars—but they require substantive additional work rather than a simple textual fix. The editor may also wish to ask the authors to verify the F1-score error in Table 7 and to clarify the 'first time' claim regarding physical tests on YOLOv8/v11/v12."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The thing to know about AdvReal is that its headline physical result—70.13% ASR on YOLOv12 in physical scenarios—is not a physical result. That number is the closed-box digital ASR from Table 3; the physical experiments in Section 4.7 use YOLOv5 as the victim detector, and Tables 10-11 contain no YOLOv12 row. The '>90% at 4m' claim also doesn't appear as angle-by-distance data anywhere. So the abstract overstates the evidence substantially.\n\nWhat the paper actually does well is assemble a joint 2D/3D adversarial training loop that mixes known ingredients—non-rigid cloth deformation, time-space mapping, relighting, ShakeDrop—into one framework, and evaluates it on a broad set of detectors. The digital results are internally consistent, the ablations show each component contributes, and the margins over AdvTexture and T-SEA are large. Code is released, which is worth credit.\n\nSoft spots beyond the abstract: the physical tables have no error bars or trial counts; the text says 'each patch was taken 111 times' but the per-cell percentages imply denominators around 37 frames, which is a big inconsistency. The confidence intervals are wide enough that some reported orderings (e.g., AdvReal-YOLOv2 at 3m vs 4m under poor light) may not be reliable. There's also no clean-clothing control, so the ASR at 4m could include natural detection failures at small scale. These issues matter because the paper's central claim is physical robustness.\n\nMy take: the digital contribution is real and the framework is worth a serious look, but the physical claims need major revision. The authors should fix the abstract, release per-trial metadata with error bars, add a non-adversarial control, and clarify the protocol (including which detector is actually the victim in physical tests).\n\nThis paper is for adversarial ML and AV perception researchers who want a stronger transferable patch baseline. It deserves peer review—the method and digital evaluation are substantive—but it should come back with the physical claims cleaned up.","headline":"The paper's headline physical claim is actually a digital closed-box result; the framework and digital evaluation are solid, but the physical claims need major revision.","tokens_in":24030,"tokens_out":2334,"would_cite":false,"duration_ms":19565,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper proposes a joint 2D–3D adversarial training framework whose printed clothing textures evade modern pedestrian detectors in physical tests, with an average attack success rate of 70.13% on YOLOv12.","keywords":["adversarial patch","physical adversarial attack","pedestrian detection","object detection","autonomous driving perception","3D rendering","non-rigid cloth deformation","attack success rate"],"falsifier":"Print the released AdvReal patch, have subjects wear it in an independently run test with a different camera, different body types, and street backgrounds not used in training, and measure the attack success rate on YOLOv12 under the paper's distance and lighting conditions; if the 70% physical ASR cannot be reproduced and the rate falls to the level of a random texture, the sim-to-real transfer claim is refuted.","tokens_in":22930,"feed_emoji":"👕","tokens_out":10882,"duration_ms":77625,"temperature":0.7,"pith_summary":"The paper seeks to establish that physical-world adversarial patches for pedestrian detectors become far more reliable when the patch is optimized jointly in ordinary 2D images and in rendered 3D scenes that mimic real street conditions. It introduces AdvReal, a training pipeline that adds stress-based cloth-wrinkle modeling, viewpoint and scale matching, and relighting to the usual patch optimization, then prints the resulting texture onto clothing. The reported result is an average attack success rate of 70.13% against YOLOv12 in physical scenarios, compared with 21.65% for T-SEA and 19.70% for AdvTexture, with success above 90% at four meters from both frontal and oblique views. If the transfer from synthetic renders to real clothing is genuine, the framework offers a practical way to stress-test the pedestrian perception modules of autonomous vehicles before deployment.","feed_headline":"Clothing patch fools YOLOv12 in 70% of physical tests","feed_subtitle":"Joint 2D-3D training with simulated wrinkles and lighting beats T-SEA and AdvTexture on real pedestrians.","key_machinery":"The load-bearing mechanism is a joint 2D–3D adversarial optimization pipeline with a realism-enhancement module. Stress-tensor estimates over garment mesh vertices (Eqs. 2–8) identify fabric regions under high tension, place spatially separated control points there, and apply noise-injected, stress-bounded deformation to generate natural wrinkles. A time-space mapping step derives camera distance, elevation, and azimuth from bounding-box sequences so the rendered human matches the background's scale and perspective, and a relighting step (Algorithm 1, Eq. 10) optimizes contrast, brightness, and bias coefficients against the structural similarity (SSIM) of the real background. Both 2D and 3D synthesized images are fed to the victim detector, and the patch is updated by the weighted detection loss $L_{total} = \\mu_1 L_{det}^{patch} + \\mu_2 L_{det}^{real} + \\mu_3 L_{tv}$ (Eq. 17). Stochastic-depth-style fusion of identity and stacked-layer outputs in the detector's residual blocks, applied to forward and backward passes, diversifies gradients during training and supports transfer to unseen detectors.","core_discovery":"The central claim is that realistic rendering, not just stronger optimization, is what makes adversarial patches survive the trip from a computer monitor to a printed garment. AdvReal trains a patch on 2D pedestrian images, then renders the same patch on a deformable 3D human mesh; stress estimates over the garment mesh place control points in high-wrinkle regions, a time-space mapping step aligns the rendered person to the perspective and scale of real background frames, and a structural-similarity-based relighting step adjusts brightness and contrast so the rendered person looks like part of the scene. Detection loss is computed on both the 2D and 3D synthesized images, with a stochastic-depth-style regularization that perturbs forward and backward passes through the detector's residual blocks to improve transferability. The authors report that this combination raises physical attack success rates on modern YOLO detectors far above existing patch methods, and that class-activation visualizations show the patch redirects the detector's attention away from the person's body.","pith_inferences":["Beyond the paper: the sim-to-real transfer is the crux, so an independent replication that varies body shape, garment fit, fabric type, and camera sensor would show whether the reported physical ASR is a property of the training method or of the specific printed samples.","Beyond the paper: the attention-hijacking behavior the paper visualizes suggests a testable defensive extension, such as training detectors to penalize attention concentrated on patch-like regions, though the paper does not evaluate such a defense.","Beyond the paper: the authors' observation that the patch works at macro scale and at low resolution implies that minor printing misalignment or laundering wear may degrade it less than high-frequency adversarial patches, a property that could be measured directly.","Beyond the paper: because the physical tests use fixed cameras and recorded video, a natural next experiment is a live test with a moving vehicle and onboard cameras to see whether the patch's robustness survives real optical pipelines and motion blur."],"forward_implications":["If the reported physical ASR transfers beyond the authors' test setup, printed AdvReal clothing provides a repeatable way to audit pedestrian detectors before deployment, since the code and textures are released.","The 70.13% physical ASR on YOLOv12 implies that one-stage modern YOLO detectors remain vulnerable to localized texture attacks, so defenses should target attention hijacking rather than pixel-level noise alone.","Because the paper reports stable ASR across distances and lighting conditions, the same patch may remain effective as a vehicle's camera approaches a pedestrian, rather than only at a fixed shooting distance.","The digital transfer across YOLOv2 through YOLOv12, Faster R-CNN, and D-DETR suggests that 3D realism training reduces the patch's overfitting to the glass-box detector's feature space."],"supporting_citations":[{"why":"Supplies the T-SEA baseline and the stochastic-depth-style regularization plus loss-weighting strategy that AdvReal adopts.","marker":"[27]"},{"why":"Provides the AdvTexture ring-cropping texture baseline against which AdvReal reports its largest physical and digital gains.","marker":"[12]"},{"why":"Shows the AdvTshirt video-based non-rigid patch mapping approach that motivates the 3D cloth deformation module.","marker":"[14]"},{"why":"Introduces TopoProj non-rigid deformation for natural-looking clothing textures, the direct predecessor extended by stress-based wrinkle modeling.","marker":"[19]"},{"why":"Argues that simple patch transforms do not capture real-world conditions, motivating the realism-enhancement pipeline and its evaluation.","marker":"[20]"},{"why":"Precedent for training 3D adversarial cloaks across poses, spatial positions, and camera angles.","marker":"[13]"},{"why":"Supplies the real traffic-scene images used as backgrounds for 3D rendering in training and testing.","marker":"[21]"},{"why":"Supplies the pedestrian images used for 2D adversarial training.","marker":"[46]"},{"why":"Provides the class-activation visualization method used to show that the patch redirects detector attention.","marker":"[28]"}],"fun_headline_variants":["Clothing patch with 3D realism fools YOLOv12 in 70% of physical trials","Adversarial garment texture hits 70% ASR on YOLOv12 in real world","YOLOv12 evaded by realistic printed patch: 70% physical success","Deformable patch trick: 70% attack rate on YOLOv12 via 2D-3D realism"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole approach depends on the computer-rendered 3D human, with simulated wrinkles and lighting, being a faithful stand-in for a real person wearing printed fabric in a real street scene; if real cloth, body motion, camera noise, or illumination differ enough from the render, the optimized textures will not transfer to the physical world.","fun_headline_variants_meta":{"raw":{"variants":["Clothing patch with 3D realism fools YOLOv12 in 70% of physical trials","Adversarial garment texture hits 70% ASR on YOLOv12 in real world","YOLOv12 evaded by realistic printed patch: 70% physical success","Deformable patch trick: 70% attack rate on YOLOv12 via 2D-3D realism"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000875,"raw_usage":{"total_tokens":3847,"prompt_tokens":1067,"completion_tokens":2780,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":683,"completion_tokens_details":{"reasoning_tokens":2678}},"tokens_in":683,"tokens_out":2780,"duration_ms":17826,"temperature":1.0,"reasoning_tokens":2678,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T15:01:42.561666+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Print the released AdvReal patch, have subjects wear it in an independently run test with a different camera, different body types, and street backgrounds not used in training, and measure the attack success rate on YOLOv12 under the paper's distance and lighting conditions; if the 70% physical ASR cannot be reproduced and the rate falls to the level of a random texture, the sim-to-real transfer claim is refuted.","supporting_citations":[{"cited_title":"Adversarial t-shirt! evading person detectors in a physical world,","cited_arxiv_id":null,"evidence_quote":"Shows the AdvTshirt video-based non-rigid patch mapping approach that motivates the 3D cloth deformation module."},{"cited_title":"T-sea: Transfer-based self-ensemble attack on object detection,","cited_arxiv_id":null,"evidence_quote":"Supplies the T-SEA baseline and the stochastic-depth-style regularization plus loss-weighting strategy that AdvReal adopts."},{"cited_title":"Adversarial texture for fooling person detectors in the physical world,","cited_arxiv_id":null,"evidence_quote":"Provides the AdvTexture ring-cropping texture baseline against which AdvReal reports its largest physical and digital gains."},{"cited_title":"Physically realizable natural-looking clothing textures evade person detectors via 3d modeling,","cited_arxiv_id":null,"evidence_quote":"Introduces TopoProj non-rigid deformation for natural-looking clothing textures, the direct predecessor extended by stress-based wrinkle modeling."},{"cited_title":"Reap: a large-scale realistic adversarial patch benchmark,","cited_arxiv_id":null,"evidence_quote":"Argues that simple patch transforms do not capture real-world conditions, motivating the realism-enhancement pipeline and its evaluation."},{"cited_title":"Learning Transferable 3D Adversarial Cloaks for Deep Trained Detectors","cited_arxiv_id":"2104.11101","evidence_quote":"Precedent for training 3D adversarial cloaks across poses, spatial positions, and camera angles."}],"review_version":1}