{"id":"9cc37d61-4746-4585-8539-3540cb023b6a","arxiv_id":"1908.09327","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"An adversarial clothing pattern, optimized to separate an individual's features across camera views, can reduce rank-1 matching from 87.9% to 27.1% and enable impersonation at 47.1% rank-1 in physical-world tests.","lead":"This paper presents a method to print adversarial patterns on clothing that make deep person re-identification systems fail to match the wearer, or match the wearer to a chosen target person. The authors demonstrate these patterns work in physical-world tests with two re-ID models, and claim the first physical-world attacks against deep re-ID.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Physical-world headline is built on an uncontrolled and thinly sampled comparison: Δrank-1 implies a 100% no-pattern baseline, yet the abstract cites 87.9%, and per-point values are multiples of 20%.","rationale":"Reading in good faith, the paper proposes a novel optimization framework for generating camera-transformable clothing patterns and provides both digital and physical evaluations. The digital experiments (Tables 2 and 3) are controlled, include generating-set and testing-set splits, and do support the general vulnerability of deep re-ID models to adversarial clothing patterns. The reader's weakest-assumption point about the overlay/degradation simulation is reasonable, but I see a more immediate load-bearing gap in the physical-world evaluation itself. Table 4's Δrank-1 column equals 100% minus attacked rank-1 at every tested point, so the reported drop is against an unstated 100% no-pattern baseline, while the abstract compares against the 87.9% PRCS rank-1 from Table 1. The paper does mention collecting images with and without the pattern, so a controlled baseline may exist, but it is not reported. The 20% granularity of all rank-1 values further indicates that the effective statistical unit is 5 identities rather than 100 queries, making per-point estimates such as 0/5 or 4/5 highly uncertain. This does not disprove the attack, and the digital results provide independent support, so a rejection is not warranted. However, the physical-world headline numbers should not be accepted at face value until the paired control data and per-identity counts are provided. I therefore retain the reader's CONDITIONAL verdict, with the additional explicit condition that the physical-world evaluation be reported with a measured no-pattern baseline and proper uncertainty quantification.","tokens_in":13318,"tokens_out":10696,"duration_ms":110196,"concrete_test":"Require the authors to release the per-identity, per-point raw matching results and to report the no-pattern control rank-1 at the same 14 test points and same 5 identities with identical gallery construction. Compute the paired rank-1 drop (no-pattern vs. pattern) with exact binomial or Clopper-Pearson confidence intervals for each point and overall, and reconcile the stated 87.9% baseline with the measured control. If the control is indeed 100% at all points, this should be stated explicitly and the abstract should use that number; if the control is materially below 87.9%, the claimed 'from 87.9% to 27.1%' framing should be revised. If per-identity variation dominates the result, the conclusion of high physical-world success should be weakened accordingly.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is the physical-world quantitative result: rank-1 drops from 87.9% to 27.1% under Evading Attack, and Impersonation Attack reaches 47.1% rank-1. In Table 4, however, Δrank-1 equals 100% minus the attacked rank-1 at every point (e.g., P2: rank-1 20%, Δrank-1 80%; P13: rank-1 80%, Δrank-1 20%). This means the reported drop is computed against a 100% no-pattern baseline, not against the 87.9% PRCS rank-1 cited in the abstract and Table 1. The paper says images were taken with/without the pattern at the 14 points, but no controlled no-pattern rank-1 values are reported, so the abstract's 'decreases from 87.9%' is an apples-to-oranges comparison with a different dataset split. Additionally, the granularity of Table 4 is telling: all 14 rank-1 values are multiples of 20%, consistent with only 5 adversary identities scored as binary successes per point rather than 100 independent queries per point. Thus the average 27.1% rests on at most 70 subject-level outcomes, and per-point estimates such as P1 0/5 and P13 4/5 have very wide confidence intervals. Without a reported paired no-pattern control and per-identity/per-query counts, the flagship physical-world effectiveness numbers are not rigorously established.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes advPattern, an optimization-based method to generate printable clothing patterns that cause deep person re-identification (re-ID) models to fail to match the wearer (Evading Attack) or to match the wearer as a chosen target person (Impersonation Attack). The pattern is optimized over a multi-position sampling set with perspective transforms, a mask to keep it decorative-looking, total variation smoothing, and a degradation model for physical-world robustness. Experiments are reported in the digital domain on two re-ID models using Market1501 and a new PRCS dataset, and in the physical world with a printed pattern worn by five adversaries at 14 positions under three cameras. The headline claims are that the rank-1 accuracy of the re-ID model for matching the adversary drops from 87.9% to 27.1% under Evading Attack, and that Impersonation Attack reaches 47.1% rank-1 and 67.9% mAP in the physical world.","tokens_in":13614,"tokens_out":4392,"duration_ms":42883,"significance":"If the physical-world claims are established, this would be an important contribution: it is, to the authors' knowledge, the first demonstration of physical-world adversarial clothing against deep re-ID, and it extends prior physical adversarial examples from classifiers to an image retrieval task. The optimization framework is clearly specified, the digital experiments are run on two different model architectures, the PRCS dataset is a useful resource, and the authors release code. The core idea of learning transformable patterns via multi-position sampling with a degradation function is sensible. However, the paper's central quantitative claims rest on the physical-world evaluation, and that evaluation currently has methodological gaps—missing no-pattern controls, an apparent mismatch between the stated and effective sample size, and internal numeric inconsistencies—that prevent the reader from verifying the headline numbers. The contribution is therefore promising but not yet rigorously supported.","major_comments":[{"comment":"The physical-world Evading Attack drop is computed against an implicit 100% no-pattern baseline: in every row of Table 4, Δrank-1 = 100% − rank-1. The paper states that images were taken 'with/without the adversarial pattern', but no per-point no-pattern rank-1 or mAP values are reported, so the reader cannot tell how much of the drop is due to the pattern and how much is due to the re-ID model's baseline error at that distance/angle. The abstract's 'decreases from 87.9% to 27.1%' compares the physical-world attacked rank-1 against the digital PRCS rank-1 of model A from Table 1, not against a paired physical-world no-pattern measurement. This apples-to-oranges comparison is load-bearing for the central claim and must be replaced with a proper paired control.","section":"Table 4, Section 5.3, Abstract"},{"comment":"The granularity of the physical-world results is inconsistent with the stated protocol. The text says '100 queries for each testing point are performed', yet all reported rank-1 values are multiples of 20% (0%, 20%, 40%, 60%, 80%), which implies only 5 binary subject-level outcomes per point. With 5 adversaries, the per-point estimates have very wide confidence intervals (e.g., a reported 0% has an upper 95% bound around 52%), and the averaged 27.1% is based on at most 70 independent identity-level outcomes. The paper must report the actual number of queries and per-identity/per-query counts, or explicitly state that the statistics are subject-level, and should provide error bars or a significance test.","section":"Section 5.3, Table 4"},{"comment":"The optimization formulations are internally inconsistent. In Eq. (8), the Evading Attack minimizes a loss plus the total variation penalty κ·TV(δ), which correctly encourages smoothness. In Eq. (9), the Impersonation Attack is written as an arg max with a positive term +κ·TV(δ); maximizing total variation is the opposite of smoothing and contradicts the stated goal of producing 'smooth and consistent patches'. In addition, the text following Eq. (8) says 'where λ and κ are hyperparameters', but λ does not appear in Eq. (8). These inaccuracies make it impossible to reproduce the exact objective used for the physical impersonation pattern.","section":"Section 4.3, Eqs. (8) and (9)"},{"comment":"The physical-world evaluation lacks the baselines and ablations needed to attribute the results to the proposed algorithm. No comparison is made to a random pattern, a plain/texture pattern, an unoptimized mask, or a pattern optimized without the multi-position/degradation components. Furthermore, physical Evading Attack is tested only with model A and physical Impersonation Attack only with model B, so there is no evidence of cross-model generalization of the physically printed pattern. Adding at least a random-pattern control and reporting both attack types on both models would substantially strengthen the claims.","section":"Section 5.3, experiment setup"}],"minor_comments":[{"comment":"The text states 'The average Δrank-1 and ΔmAP are 62.2% and 61.1%', but Table 4 shows the average Δrank-1 as 74.3%; only ΔmAP (61.1%) matches. Please correct the inconsistency.","section":"Section 5.3, text vs. Table 4"},{"comment":"The sentences 'For Evading Attack, the average rank-1 accuracy drops to 11.1% in 9 of 14 positions' and 'The rank-1 accuracy for matching the adversary as the target person is 56.4% in 11 of 14 positions' are unclear: are these averages over the selected subsets, and how do they relate to the overall averages of 27.1% and 47.1%? Please clarify and report the per-subset values directly.","section":"Section 5.3, text"},{"comment":"The symbol D(δ) in the threat-model objective is introduced as a measure of the 'reality' of the pattern, but its definition is never given and it does not appear in the implemented optimization objectives (Eqs. (3)–(9)). Either define D(δ) and show how it is incorporated, or remove it from the formal problem statement to avoid a dangling term.","section":"Section 4.1, Eq. (1)"},{"comment":"The text mentions that NPS was introduced but 'hard to balance' and then replaced by a color-interval constraint P; this is fine, but the sentence structure reads as if NPS were still part of the final objective. Please rephrase to state clearly that NPS is not used in the final formulation.","section":"Section 4.3"},{"comment":"The similarity-score column (ss) is not defined; the reader must infer that it is an average pairwise similarity. Please provide the exact formula or definition, as it is later used to interpret the impact of attacks.","section":"Section 5.1, Table 1"},{"comment":"There is a typo in 'the ﬁled of cameras views'; and the left-top point is said to be omitted but its position is not indicated in the figure. Also, the distance range in the text (5m–10m) does not match the example coordinates in Figure 5 (e.g., P1 at 4.39m). Please align these details.","section":"Figure 5 and Section 5.3"},{"comment":"The total variation term appears to be an L2 TV; if the intent is to encourage piecewise-smooth patches, an L1 TV is often used. The choice is not motivated. This is a minor presentation issue but should be clarified.","section":"Section 4.3, Eq. (7)"}],"recommendation":"major_revision","confidential_remarks":"The manuscript carries a footnote stating it was accepted by IEEE ICCV 2019. If this is a journal submission of that work, the review should focus on whether the arXiv version meets the journal's standard; the footnote is unusual in a journal submission and may need to be removed or converted into a prior-publication disclosure. The main technical concern is that the physical-world evaluation, which carries the headline claims, does not currently provide a valid comparison baseline or an honest sample-size description; these issues are correctable with additional experiments and reporting, so I do not recommend rejection on the current evidence."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe headline: this is the first physical-world adversarial attack on person re-ID, and the idea of generating a clothing pattern that transfers across camera views is a real contribution. But the flagship physical-world numbers are not as clean as the abstract implies; Table 4's Δrank-1 is computed against a 100% no-pattern baseline, not the 87.9% cited as the starting point, and the rank-1 granularity suggests tiny per-point samples.\n\nThe paper does several things well. The cross-camera separation objective in Eqs. (3)-(6) makes sense for re-ID, the multi-position sampling for viewpoint scalability is sensible, and the digital attacks show a clear drop on held-out testing sets. The authors also built a new dataset, PRCS, and released code. That is reproducible progress.\n\nThe soft spots are concentrated in the physical-world evaluation, and the stress-test note lands. In Table 4, at every one of the 14 points, Δrank-1 = 100% − rank-1. That means no-pattern rank-1 is implicitly 100% at each point, not the 87.9% from Table 1 quoted in the abstract. The claim \"decreases from 87.9% to 27.1%\" mixes a global baseline with per-position attacked numbers, which is apples-to-oranges. Also, all rank-1 values are multiples of 20%, consistent with 5 binary identity outcomes per point, not the stated 100 queries. P1 has 0.0% and P13 80%: with n=5 the confidence intervals are huge. No error bars or significance tests appear. This is a load-bearing issue because the physical-world success is the paper's main claim.\n\nA few smaller issues: there are no baselines with random or non-adversarial decorative patterns, so part of the drop could come from a conspicuous outfit. Only one model is used per physical attack. The \"position-irrelevant\" claim is contradicted by four points with 0% evading success. Hyperparameters like α, β, λ1, λ2, and the color interval P are not fully specified, though code is available. The impersonation results are partly circular (you optimize similarity to the target), but the digital held-out transfer and the physical-world transfer to unseen positions are the non-circular parts.\n\nWho should read this: anyone working on adversarial robustness for retrieval or surveillance systems. It deserves a serious referee; the reviewer should focus on the physical-world statistics and demand a paired no-pattern control at the same 14 positions, plus per-identity counts. With that fixed, it could be a solid paper. As is, I'd treat the abstract's headline numbers as unverified.","headline":"First physical-world re-ID attack with a genuinely new cross-camera pattern idea, but the headline physical-world numbers rest on a baseline inconsistency and tiny effective samples.","tokens_in":14147,"tokens_out":3252,"would_cite":true,"duration_ms":28874,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Adversarial clothing patterns can hide or impersonate a person across surveillance cameras.","keywords":["adversarial examples","person re-identification","physical-world attack","advPattern","evasion attack","impersonation attack","clothing patterns","cross-camera matching"],"falsifier":"A controlled experiment in which the same pattern is printed by two different printers and photographed under controlled lighting at the same 14 positions: if the attack success rate changes materially between printers or lighting conditions, the degradation function $\\phi(\\cdot)$ does not capture the physical factors the paper claims it models.","tokens_in":13126,"feed_emoji":"👕","tokens_out":5683,"duration_ms":52330,"temperature":0.7,"pith_summary":"This paper tries to establish that deep person re-identification (re-ID) systems, which match people across camera views, can be attacked by wearing a printed adversarial pattern on clothing. It proposes advPattern, an algorithm that optimizes a decorative-looking patch so the wearer's image features move apart across cameras in an Evading Attack, or move toward a chosen target's features in an Impersonation Attack. The reported physical-world results show re-ID rank-1 matching of the wearer falling from 87.9% to 27.1% under evasion, with impersonation achieving 47.1% rank-1 accuracy and 67.9% mAP. If correct, this means a person could evade or hijack person search in surveillance without altering any stored image.","feed_headline":"Printed clothing pattern hides or fakes identity in re-ID cameras","feed_subtitle":"Wearing the pattern cuts re-ID rank-1 from 87.9% to 27.1%; impersonation hits 47.1%.","key_machinery":"The key machinery is advPattern, an iterative optimization that adjusts a clothing pattern by minimizing or maximizing similarity scores over image pairs. For each image the pattern is perspectively transformed with $T_i(\\delta)$ and overlaid via $o(x_i, T_i(\\delta))$; a multi-position sampling strategy augments the generating set with varied distances and angles, and a degradation function $\\phi(\\cdot)$ randomly changes brightness or blurs images to simulate physical dynamics. A total-variation term smooths the pattern, a mask $M_x$ shapes it like a decorative logo, and the search space is constrained to a printable color interval $P$. This machinery converts an appearance change into a controlled shift of the person's location in the re-ID feature space, which is what makes evasion and impersonation possible.","core_discovery":"The central discovery is that adversarial patterns printed on clothing can transfer across camera views and fool deep re-ID models in the physical world. Concretely, the paper claims that a wearer's rank-1 matching accuracy drops from 87.9% to 27.1% under an Evading Attack, and that under an Impersonation Attack the wearer is matched to a chosen target person with 47.1% rank-1 accuracy and 67.9% mAP. The attack is achieved without modifying any stored image: the adversary only wears the pattern, and the optimization makes the re-ID feature extractor pull same-camera images together while pushing cross-camera images apart (evasion), or pull the wearer's features toward the target's features (impersonation). The paper further claims that both a siamese network and a classification-based embedding network are vulnerable.","pith_inferences":["If the pattern's effectiveness transfers to unseen re-ID architectures, a practical countermeasure would be to train re-ID models with adversarial-pattern augmentation so identity features ignore clothing texture; this is a testable extension the paper does not explore.","Because four of the fourteen physical test positions showed 0% Evading Attack success, one extension is to sample more positions and lighting conditions during optimization; the paper's own data implies the current degradation model under-covers some real-world conditions.","Cross-dataset impersonation (target from Market1501) succeeded less than same-dataset impersonation, suggesting that domain gap, not just pattern optimization, limits physical impersonation; measuring how much of the gap is due to image style could guide improved transfer.","The white-box assumption is strong; a natural next test is whether patterns optimized on one model transfer to a black-box re-ID system, since adversarial transferability has been shown in classification."],"forward_implications":["If the central claim holds, a person wearing the generated pattern can expect to evade cross-camera matching in roughly three of every four physical-world queries (average rank-1 down to 27.1%).","The same pattern can make a re-ID system name a chosen target for the wearer almost half the time (47.1% rank-1) under impersonation, meaning targeted misidentification is physically realizable.","Attack effectiveness varies with position: at some camera distances and angles the pattern gives 0% matching success, while at others it leaves evasion rank-1 at 60-80%, so the attack is position-dependent.","Both siamese-style and classification-style deep re-ID models are vulnerable, which suggests the vulnerability is not an artifact of one architecture."],"supporting_citations":[{"why":"Introduces adversarial examples, the phenomenon this paper applies to re-ID.","marker":"[28]"},{"why":"Provides the fast gradient sign method baseline for generating perturbations.","marker":"[10]"},{"why":"Demonstrates printed adversarial examples can fool classifiers when photographed.","marker":"[15]"},{"why":"Shows physical adversarial accessories can attack face recognition, serving as a template for wearable patterns.","marker":"[25]"},{"why":"Shows physical-world adversarial road signs survive environmental conditions, motivating degradation modeling.","marker":"[8]"},{"why":"Shows 3D-printed adversarial objects fool classifiers over viewpoints, supporting multi-position sampling.","marker":"[2]"},{"why":"Provides model A, the siamese re-ID network the attacks are evaluated against.","marker":"[37]"},{"why":"Provides model B, the classification-based re-ID network the attacks are evaluated against.","marker":"[36]"},{"why":"Documents camera-specific style variation that motivates cross-camera transformable patterns.","marker":"[38]"}],"fun_headline_variants":["Printed pattern drops re-ID rank-1 from 87.9% to 27.1%","Wearable adversarial pattern evades re-ID, impersonates targets","Clothing pattern fools deep re-ID: match rate falls to 27%","Adversarial pattern on clothes hides identity from re-ID cameras"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The attack works only if overlaying the digitally generated pattern and applying the brightness/blur degradation function faithfully reproduces how the printed fabric appears in real camera views across distances and angles.","fun_headline_variants_meta":{"raw":{"variants":["Printed pattern drops re-ID rank-1 from 87.9% to 27.1%","Wearable adversarial pattern evades re-ID, impersonates targets","Clothing pattern fools deep re-ID: match rate falls to 27%","Adversarial pattern on clothes hides identity from re-ID cameras"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000343,"raw_usage":{"total_tokens":1921,"prompt_tokens":1016,"completion_tokens":905,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":632,"completion_tokens_details":{"reasoning_tokens":822}},"tokens_in":632,"tokens_out":905,"duration_ms":7971,"temperature":1.0,"reasoning_tokens":822,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T11:14:51.464828+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A controlled experiment in which the same pattern is printed by two different printers and photographed under controlled lighting at the same 14 positions: if the attack success rate changes materially between printers or lighting conditions, the degradation function $\\phi(\\cdot)$ does not capture the physical factors the paper claims it models.","supporting_citations":[{"cited_title":"Accessorize to a crime: Real and stealthy attacks on state-of-the-art face recognition","cited_arxiv_id":null,"evidence_quote":"Shows physical adversarial accessories can attack face recognition, serving as a template for wearable patterns."},{"cited_title":"Robust physical-world attacks on deep learning visual classiﬁcation","cited_arxiv_id":null,"evidence_quote":"Shows physical-world adversarial road signs survive environmental conditions, motivating degradation modeling."},{"cited_title":"A discrimi- natively learned cnn embedding for person reidentiﬁcation","cited_arxiv_id":null,"evidence_quote":"Provides model A, the siamese re-ID network the attacks are evaluated against."},{"cited_title":"Camera style adaptation for person re- identiﬁcation","cited_arxiv_id":null,"evidence_quote":"Documents camera-specific style variation that motivates cross-camera transformable patterns."}],"review_version":1}