REVIEW 2 major objections 6 minor 76 references
Modality Unified Attack for Omni-Modality Person Re-Identification
T0 review · 2 major / 6 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read The paper claims that three modality-specific adversarial generators, trained together on a single multi-modality surrogate model with metric disruption losses, can produce transferable perturbations that degrade single-, cross-, and…
desk verdict First genuine omni-modality re-id attack with broad black-box evaluation; the CMSD transfer mechanism is asserted rather than proven, but the paper deserves review. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The machinery is a multi-modality surrogate model with three modality-specific subnetworks, plus three trained adversarial generators and one discriminator per modality. The decisive object is Cross Modality Simulated Disruption: feeding an RGB image into the NI subnetwork and an NI image into the RGB subnetwork to fabricate modality-invariant features, then using Euclidean-distance disruption on those simulated features so the generators attack the shared embedding space that cross-modality models rely on. Multi Modality Collaborative Disruption plays the complementary role of pushing each adversarial feature away from the feature centers of all three modalities, so that no useful complementary information survives the fusion step.
What would settle it
Measure the distance between simulated cross-modality features $S_R(x_N)$ and $S_N(x_R)$ from the surrogate and the embeddings a real cross-modality model produces for the same images; if the simulated and real feature distributions are far apart, or if adversarial examples trained only to disrupt the simulated features fail to raise retrieval rank on cross-modality models, the central transfer mechanism is not doing the claimed work.
Extended reading notes
Core claim
The paper sets out to show that a unified black-box attack on person re-identification is possible: instead of designing one attack per model family, three modality-specific generators for RGB, near-infrared, and thermal infrared images are trained together on a surrogate multi-modality model by disrupting intermediate features before fusion. The proposed Modality Unified Attack combines Metric Disruption with two new losses: Cross Modality Simulated Disruption, which feeds images through non-corresponding modality subnetworks to approximate cross-modality embedding spaces, and Multi Modality Collaborative Disruption, which pushes adversarial features away from all three modality feature centers. In cross-dataset and cross-model black-box experiments, the same generators produce mean mAP Drop Rates of 55.9% on eight single-modality models, 24.4% and 49.0% on the two cross-modality retrieval directions, and 62.7% on two multi-modality models, and remain effective against JPEG compression and input randomization defenses.
Load-bearing premise
The attack transfers only if feeding an image into the wrong-modality subnetwork of the surrogate produces features close to those of a real cross-modality re-id model; if that simulation is unrepresentative, the cross-modality attack gains would not generalize.
Editorial extensions
If this is right
- One set of per-modality generators replaces the need to know the target model type in a black-box re-id attack.
- Cross-modality retrieval in both directions, RGB-to-NI and NI-to-RGB, is degraded by the same generators that attack single- and multi-modality models.
- The Cross Modality Simulated Disruption loss is the main driver of cross-modality transfer: adding it raises the RGB-to-NI mean mAP Drop Rate from 10.4% to 24.1%.
- Multi-modality feature fusion does not by itself neutralize the attack: the two tested fused models drop by 62.7% mean mAP Drop Rate.
- Basic input defenses weaken but do not stop the attack, with mean mAP Drop Rates still at 41.0%, 18.1%, 39.5%, and 36.7% under JPEG and randomization defenses.
Reading between the lines
- The paper does not compare its simulated cross-modality features directly with embeddings from real cross-modality models; a feature-space similarity measurement would be the cleanest way to test whether the transfer works for the reason claimed.
- The same surrogate-subnetwork trick could be applied to other tasks with per-modality encoders, such as visible-infrared detection, remote sensing, or multimodal retrieval, whenever an attacker wants one perturbation set to work across single- and cross-modal systems.
- Because all experiments use one surrogate backbone and one multi-modality dataset, the unified claim would be stronger if the generators were retrained on a second multi-modality backbone and evaluated on a second multi-modality benchmark; that remains untested.
- The Multi Modality Collaborative Disruption constraint could be reused as a robustness diagnostic during adversarial training, since it directly targets the complementary information that fusion models depend on.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a generative adversarial attack, Modality Unified Attack (MUA), for person re-identification models across single-, cross-, and multi-modality settings. MUA trains three modality-specific generators (RGB, NI, TI) against a multi-modality surrogate model (TOP), using a metric disruption loss, a Cross Modality Simulated Disruption (CMSD) loss that pushes apart features obtained by cross-feeding images into non-corresponding subnetworks, and a Multi Modality Collaborative Disruption (MMCD) loss that pushes adversarial features away from feature centers of all modalities. The generators are evaluated for black-box transfer on 14 models spanning the three settings, with reported mean mAP drop rates of 55.9%, 24.4%, 49.0%, and 62.7%. Ablations show that each loss component contributes, with CMSD being the key component for cross-modality transfer.
Significance. If the reported transfer rates are reproducible, the paper is a valuable first step toward understanding the security of multi-modality re-id systems under black-box attacks. The choice of a multi-modality surrogate, intermediate-feature disruption before fusion, and evaluation across diverse architectures and datasets are all strengths; the ablation study is informative, and the cross-dataset/cross-model protocol is appropriate for single- and cross-modality targets. However, the central simulation assumption underlying CMSD is not validated, and the multi-modality setting lacks any comparison baseline, so the headline claims are not yet fully supported.
major comments (2)
- [Section III-D, Eq. (9) and Table II] The second equality in Eq. (9), F^{N'}_R = S_R(x'_R), is inconsistent with the text and with Eq. (8); it should read F^{N'}_R = S_N(x'_R). Because the CMSD loss (Eq. 10) directly uses F^{N'}_R, the printed equation makes the implementation ambiguous. Beyond the typo, the core premise of CMSD is asserted rather than validated: the authors assume that S_N(x_R), where S_N is trained only on NI images, is a faithful stand-in for the modality-invariant embedding of a cross-modality model. Since Table II attributes the largest cross-modality gains to CMSD (R-N mDR from 10.4 to 24.4; N-R from 41.4 to 49.0), this assumption is load-bearing. I request a direct feature-space check (e.g., distance between S_N(x_R) and embeddings from CAJ/DDAG/PMT on LLCM) or an ablation that replaces the cross-input with a fixed random projection, to show that the gain is not simply due to out-of-distribution sensitivity.
- [Section IV-A, Table I] There is no attack baseline for the multi-modality retrieval setting (R'N'T'-RNT). The 62.7% mDR is meaningful relative to 'None', but the claim that MUA is the first omni-modality attack does not by itself establish that the multi-modality result is competitive or non-trivial. Please add at least one baseline, even a simple one, such as training three independent single-modality generators with the MD loss on the same surrogate and applying them to the three modalities, or adapting an existing cross-modality attack (e.g., FSAM) to TI images. If none can be adapted, explicitly justify why the 62.7% mDR is meaningful without a comparison.
minor comments (6)
- [Section III-E, Eq. (11)] The 'individual feature centers' F^c_m are not defined. Specify how they are computed (e.g., exponential moving average over the training set or per-batch mean) and over which data split.
- [Section IV-C, Fig. 5(a)] The text states that 'high lambda1 and lambda2 values lead to notable mDR declines,' yet the chosen values are 50. Please clarify the parameter axis and where the optimum lies; as written, the sentence is ambiguous.
- [Section IV-C] In the 'Effectiveness of MD' paragraph, the text reports 58.2% mDR for the multi-modality setting, while Table II reports 58.6%; please correct the discrepancy.
- [Section IV-B] The table would be clearer if the caption or a note indicated which methods are designed for which settings; currently the grouping into 'Single-modality Attack', 'Cross-modality Attack', and 'Omni-modality Attack' is not explicitly explained.
- [Section I] The phrase 'complementary feature fusion' appears twice in close succession; consider rewording one instance for clarity.
- [Figure 2] The notation F^N_R and F^{N'}_R is not defined in the caption; please define it in the caption or in the text.
Circularity Check
No significant circularity: the reported mDR values are measured on held-out black-box models, not on the surrogate used to train the generators.
full rationale
The derivation chain is not circular. The three generators are trained only against the surrogate multi-modality model TOP using feature-space losses (Eqs. 7, 10, 11) on the surrogate's own subnetworks and feature centers, while the headline results in Table I are computed on held-out black-box single-, cross-, and multi-modality re-id models that do not participate in the generator training objective. No equation reduces the evaluation metric to the training objective by construction: L_MD, L_CMSD, and L_MMCD maximize Euclidean separation in the surrogate's intermediate feature space, whereas mDR measures retrieval-rank degradation on distinct target models. The hyperparameters lambda1, lambda2, lambda3 are reported as selected choices after an ablation, which is standard model selection rather than a fitted prediction. The CMSD premise that feeding images into non-corresponding subnetworks mimics cross-modality embedding spaces is asserted rather than directly validated, and Eq. 9 contains an apparent typo (S_R(x'_R) should likely read S_N(x'_R)), but these are correctness and reproducibility concerns, not circularity: the claim is empirically testable and is in fact tested against external black-box models. The self-citations (e.g., ref. [39] in the related-work enumeration) are not load-bearing for the central derivation. Therefore the paper's core attack-transfer claim has independent empirical content and no significant circularity is present.
Assumptions & free parameters
free parameters (2)
- balance weights lambda1, lambda2, lambda3 =
50, 50, 10
- perturbation bound epsilon =
8/255
assumptions (4)
- domain assumption The features extracted by a multi-modality surrogate's modality-specific subnetworks, before fusion, are representative of the embedding spaces used by single-modality re-id models.
- domain assumption Feeding an image into a non-corresponding modality subnetwork (e.g., RGB into the NI subnetwork) produces a distribution close to a cross-modality re-id model's modality-invariant embedding space.
- domain assumption The feature centers F^c_m computed from benign multi-modality features are stable enough to represent the individual feature spaces for all three modalities.
- standard math Background existence of adversarial vulnerability of deep metric-learning re-id models and transferability of GAN-based attacks.
Cite this review
Pith. "Pith review of Modality Unified Attack for Omni-Modality Person Re-Identification." pith.science (2026). https://pith.science/paper/REBWXQHY
@misc{pith2026250112761,
author = {Pith},
title = {Pith review of: Modality Unified Attack for Omni-Modality Person Re-Identification},
year = {2026},
howpublished = {\url{https://pith.science/paper/REBWXQHY}},
note = {Machine review of arXiv:2501.12761}
}
read the original abstract
Deep learning based person re-identification (re-id) models have been widely employed in surveillance systems. Recent studies have demonstrated that black-box single-modality and cross-modality re-id models are vulnerable to adversarial examples (AEs), leaving the robustness of multi-modality re-id models unexplored. Due to the lack of knowledge about the specific type of model deployed in the target black-box surveillance system, we aim to generate modality unified AEs for omni-modality (single-, cross- and multi-modality) re-id models. Specifically, we propose a novel Modality Unified Attack method to train modality-specific adversarial generators to generate AEs that effectively attack different omni-modality models. A multi-modality model is adopted as the surrogate model, wherein the features of each modality are perturbed by metric disruption loss before fusion. To collapse the common features of omni-modality models, Cross Modality Simulated Disruption approach is introduced to mimic the cross-modality feature embeddings by intentionally feeding images to non-corresponding modality-specific subnetworks of the surrogate model. Moreover, Multi Modality Collaborative Disruption strategy is devised to facilitate the attacker to comprehensively corrupt the informative content of person images by leveraging a multi modality feature collaborative metric disruption loss. Extensive experiments show that our MUA method can effectively attack the omni-modality re-id models, achieving 55.9%, 24.4%, 49.0% and 62.7% mean mAP Drop Rate, respectively.
Figures
Figures from the paper (4 more)
Reference graph
Works this paper leans on
-
[1]
Deep learning for person re-identification: A survey and outlook,
M. Ye, J. Shen, G. Lin, T. Xiang, L. Shao, and S. C. Hoi, “Deep learning for person re-identification: A survey and outlook,” IEEE Trans. Pattern Anal. Mach. Intell. , vol. 44, no. 6, pp. 2872–2893, 2021
work page 2021
-
[2]
Person re-identification: Past, present and future,
L. Zheng, Y . Yang, and A. G. Hauptmann, “Person re-identification: Past, present and future,” arXiv preprint arXiv:1610.02984 , 2016
arXiv 2016
-
[3]
Region generation and assessment network for occluded person re- identification,
S. He, W. Chen, K. Wang, H. Luo, F. Wang, W. Jiang, and H. Ding, “Region generation and assessment network for occluded person re- identification,” IEEE Trans. Inf. Forensics Secur , vol. 19, pp. 120–132, 2024
work page 2024
-
[4]
Pseudo-label noise prevention, suppression and softening for unsupervised person re- identification,
H. Wang, M. Yang, J. Liu, and W.-S. Zheng, “Pseudo-label noise prevention, suppression and softening for unsupervised person re- identification,” IEEE Trans. Inf. Forensics Secur, vol. 18, pp. 3222–3237, 2023
work page 2023
-
[5]
Person re-identification with hierarchical discriminative spatial aggregation,
M. Zhang, Y . Xiao, F. Xiong, S. Li, Z. Cao, Z. Fang, and J. T. Zhou, “Person re-identification with hierarchical discriminative spatial aggregation,” IEEE Trans. Inf. Forensics Secur , vol. 17, pp. 516–530, 2022
work page 2022
-
[6]
Partial person re-identification,
W.-S. Zheng, X. Li, T. Xiang, S. Liao, J. Lai, and S. Gong, “Partial person re-identification,” in Int. Conf. Comput. Vis. , 2015, pp. 4678– 4686
work page 2015
-
[7]
Weakly super- vised tracklet association learning with video labels for person re- identification,
M. Liu, Y . Bian, Q. Liu, X. Wang, and Y . Wang, “Weakly super- vised tracklet association learning with video labels for person re- identification,” IEEE Trans. Pattern Anal. Mach. Intell. , vol. 46, no. 5, pp. 3595–3607, 2024
work page 2024
-
[8]
Y . Sun, L. Zheng, Y . Yang, Q. Tian, and S. Wang, “Beyond part models: Person retrieval with refined part pooling (and a strong convolutional baseline),” in Eur. Conf. Comput. Vis., 2018, pp. 480–496
work page 2018
Show all 76 references
-
[9]
Nas-ped: Neural architecture search for pedestrian detection,
Y . Tang, M. Liu, B. Li, Y . Wang, and W. Ouyang, “Nas-ped: Neural architecture search for pedestrian detection,” IEEE Transactions on Pattern Analysis and Machine Intelligence , 2024
2024
-
[10]
Spatial-temporal person re- identification,
G. Wang, J. Lai, P. Huang, and X. Xie, “Spatial-temporal person re- identification,” in AAAI Conf. Artif. Intell. , vol. 33, no. 01, 2019, pp. 8933–8940
2019
-
[11]
Diverse part dis- covery: Occluded person re-identification with part-aware transformer,
Y . Li, J. He, T. Zhang, X. Liu, Y . Zhang, and F. Wu, “Diverse part dis- covery: Occluded person re-identification with part-aware transformer,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2021, pp. 2898–2907
2021
-
[12]
Transreid: Transformer-based object re-identification,
S. He, H. Luo, P. Wang, F. Wang, H. Li, and W. Jiang, “Transreid: Transformer-based object re-identification,” in Int. Conf. Comput. Vis. , 2021, pp. 15 013–15 022
2021
-
[13]
Occlusion-aware feature recover model for occluded person re-identification,
Y . Bian, M. Liu, X. Wang, Y . Tang, and Y . Wang, “Occlusion-aware feature recover model for occluded person re-identification,” IEEE Trans. Multimedia, vol. 26, pp. 5284–5295, 2024
2024
-
[14]
A two-stage noise-tolerant paradigm for label corrupted person re- identification,
M. Liu, F. Wang, X. Wang, Y . Wang, and A. K. Roy-Chowdhury, “A two-stage noise-tolerant paradigm for label corrupted person re- identification,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 46, no. 7, pp. 4944–4956, 2024
2024
-
[15]
Visible-infrared person re-identification via homogeneous augmented tri-modal learning,
M. Ye, J. Shen, and L. Shao, “Visible-infrared person re-identification via homogeneous augmented tri-modal learning,” IEEE Trans. Inf. Forensics Secur, vol. 16, pp. 728–739, 2020
2020
-
[16]
Rgb-infrared cross- modality person re-identification,
A. Wu, W.-S. Zheng, H.-X. Yu, S. Gong, and J. Lai, “Rgb-infrared cross- modality person re-identification,” in Int. Conf. Comput. Vis. , 2017, pp. 5380–5389
2017
-
[17]
Dual-semantic consistency learning for visible-infrared person re-identification,
Y . Zhang, Y . Kang, S. Zhao, and J. Shen, “Dual-semantic consistency learning for visible-infrared person re-identification,” IEEE Trans. Inf. Forensics Secur, vol. 18, pp. 1554–1565, 2023
2023
-
[18]
Cross-modality paired-images generation for rgb-infrared person re-identification,
G.-A. Wang, T. Zhang, Y . Yang, J. Cheng, J. Chang, X. Liang, and Z.-G. Hou, “Cross-modality paired-images generation for rgb-infrared person re-identification,” in AAAI, vol. 34, no. 07, 2020, pp. 12 144–12 151
2020
-
[19]
Infrared-visible cross-modal person re-identification with an x modality,
D. Li, X. Wei, X. Hong, and Y . Gong, “Infrared-visible cross-modal person re-identification with an x modality,” in AAAI Conf. Artif. Intell., vol. 34, no. 04, 2020, pp. 4610–4617
2020
-
[20]
Towards grand unified representation learning for unsupervised visible-infrared person re-identification,
B. Yang, J. Chen, and M. Ye, “Towards grand unified representation learning for unsupervised visible-infrared person re-identification,” in Int. Conf. Comput. Vis. , 2023, pp. 11 069–11 079
2023
-
[21]
Robust multi-modality person re-identification,
A. Zheng, Z. Wang, Z. Chen, C. Li, and J. Tang, “Robust multi-modality person re-identification,” in AAAI Conf. Artif. Intell., vol. 35, no. 4, 2021, pp. 3529–3537
2021
-
[22]
Interact, embed, and enlarge: Boosting modality-specific representations for multi-modal person re-identification,
Z. Wang, C. Li, A. Zheng, R. He, and J. Tang, “Interact, embed, and enlarge: Boosting modality-specific representations for multi-modal person re-identification,” in AAAI Conf. Artif. Intell., vol. 36, no. 3, 2022, pp. 2633–2641
2022
-
[23]
Magic tokens: Select diverse tokens for multi-modal object re-identification,
P. Zhang, Y . Wang, Y . Liu, Z. Tu, and H. Lu, “Magic tokens: Select diverse tokens for multi-modal object re-identification,” arXiv preprint arXiv:2403.10254, 2024
2024 arXiv
-
[24]
Top-reid: Multi- spectral object re-identification with token permutation,
Y . Wang, X. Liu, P. Zhang, H. Lu, Z. Tu, and H. Lu, “Top-reid: Multi- spectral object re-identification with token permutation,” in AAAI Conf. Artif. Intell., vol. 38, no. 6, 2024, pp. 5758–5766
2024
-
[25]
Explaining and harnessing adversarial examples,
I. J. Goodfellow, J. Shlens, and C. Szegedy, “Explaining and harnessing adversarial examples,” arXiv preprint arXiv:1412.6572 , 2014
2014 arXiv
-
[26]
Intriguing properties of neural networks,
C. Szegedy, W. Zaremba, I. Sutskever, J. Bruna, D. Erhan, I. Goodfellow, and R. Fergus, “Intriguing properties of neural networks,” in Int. Conf. Learn. Represent., 2014
2014
-
[27]
Towards deep learning models resistant to adversarial attacks,
A. Madry, A. Makelov, L. Schmidt, D. Tsipras, and A. Vladu, “Towards deep learning models resistant to adversarial attacks,” inInt. Conf. Learn. Represent., 2018. 11
2018
-
[28]
U-turn: Crafting adversarial queries with opposite-direction features,
Z. Zheng, L. Zheng, Y . Yang, and F. Wu, “U-turn: Crafting adversarial queries with opposite-direction features,” Int. J. Comput. Vis. , vol. 131, no. 4, pp. 835–854, 2023
2023
-
[29]
Adversarial metric attack and defense for person re-identification,
S. Bai, Y . Li, Y . Zhou, Q. Li, and P. H. Torr, “Adversarial metric attack and defense for person re-identification,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 43, no. 6, pp. 2119–2126, 2020
2020
-
[30]
Vulnerability of person re-identification models to metric adversarial attacks,
Q. Bouniot, R. Audigier, and A. Loesch, “Vulnerability of person re-identification models to metric adversarial attacks,” in IEEE Conf. Comput. Vis. Pattern Recog. Worksh. , 2020, pp. 794–795
2020
-
[31]
advpattern: Physical-world attacks on deep person re-identification via adversarially transformable patterns,
Z. Wang, S. Zheng, M. Song, Q. Wang, A. Rahimpour, and H. Qi, “advpattern: Physical-world attacks on deep person re-identification via adversarially transformable patterns,” in Int. Conf. Comput. Vis. , 2019, pp. 8341–8350
2019
-
[32]
Cross-modality perturbation synergy attack for person re-identification,
Y . Gong et al., “Cross-modality perturbation synergy attack for person re-identification,” in Adv. Neural Inform. Process. Syst. , 2024
2024
-
[33]
Decoupled feature-based mix- ture of experts for multi-modal object re-identification,
Y . Wang, Y . Liu, A. Zheng, and P. Zhang, “Decoupled feature-based mix- ture of experts for multi-modal object re-identification,” arXiv preprint arXiv:2412.10650, 2024
2024 arXiv
-
[34]
Towards robust person re-identification by defending against universal attackers,
F. Yang, J. Weng, Z. Zhong, H. Liu, Z. Wang, Z. Luo, D. Cao, S. Li, S. Satoh, and N. Sebe, “Towards robust person re-identification by defending against universal attackers,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 45, no. 4, pp. 5218–5235, 2022
2022
-
[35]
Learning to attack real-world models for person re-identification via virtual-guided meta-learning,
F. Yang, Z. Zhong, H. Liu, Z. Wang, Z. Luo, S. Li, N. Sebe, and S. Satoh, “Learning to attack real-world models for person re-identification via virtual-guided meta-learning,” in AAAI Conf. Artif. Intell., vol. 35, no. 4, 2021, pp. 3128–3135
2021
-
[36]
Beyond universal person re-identification attack,
W. Ding, X. Wei, R. Ji, X. Hong, Q. Tian, and Y . Gong, “Beyond universal person re-identification attack,” IEEE Trans. Inf. Forensics Secur, vol. 16, pp. 3442–3455, 2021
2021
-
[37]
Transferable, control- lable, and inconspicuous adversarial attacks on person re-identification with deep mis-ranking,
H. Wang, G. Wang, Y . Li, D. Zhang, and L. Lin, “Transferable, control- lable, and inconspicuous adversarial attacks on person re-identification with deep mis-ranking,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2020, pp. 342–351
2020
-
[38]
Meta generative attack on person reidentification,
A. Subramanyam, “Meta generative attack on person reidentification,” IEEE Trans. Circuit Syst. Video Technol., vol. 33, no. 8, pp. 4429–4434, 2023
2023
-
[39]
Learning to learn transferable generative attack for person re-identification,
Y . Bian, M. Liu, X. Wang, Y . Ma, and Y . Wang, “Learning to learn transferable generative attack for person re-identification,”arXiv preprint arXiv:2409.04208, 2024
2024 arXiv
-
[40]
Pose-guided visible part matching for occluded person reid,
S. Gao, J. Wang, H. Lu, and Z. Liu, “Pose-guided visible part matching for occluded person reid,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2020, pp. 11 744–11 752
2020
-
[41]
Pose transferrable person re-identification,
J. Liu, B. Ni, Y . Yan, P. Zhou, S. Cheng, and J. Hu, “Pose transferrable person re-identification,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2018, pp. 4099–4108
2018
-
[42]
Illumination-invariant person re-identification,
Y . Huang, Z.-J. Zha, X. Fu, and W. Zhang, “Illumination-invariant person re-identification,” in ACM Int. Conf. Multimedia , 2019, pp. 365–373
2019
-
[43]
Learning person re-identification models from videos with weak supervision,
X. Wang, M. Liu, D. S. Raychaudhuri, S. Paul, Y . Wang, and A. K. Roy- Chowdhury, “Learning person re-identification models from videos with weak supervision,” IEEE Trans. Image Process., vol. 30, pp. 3017–3028, 2021
2021
-
[44]
Exploiting global camera network constraints for unsupervised video person re-identification,
X. Wang, R. Panda, M. Liu, Y . Wang, and A. K. Roy-Chowdhury, “Exploiting global camera network constraints for unsupervised video person re-identification,” IEEE Trans. Circuit Syst. Video Technol. , vol. 31, no. 10, pp. 4020–4030, 2021
2021
-
[45]
Feature-level adversarial attacks and ranking disruption for visible-infrared person re- identification,
X. Yang, H. Liu, D. Cheng, N. Wang, and X. Gao, “Feature-level adversarial attacks and ranking disruption for visible-infrared person re- identification,” in Adv. Neural Inform. Process. Syst. , 2024
2024
-
[46]
Cross-modal learning with adversarial samples,
C. Li, S. Gao, C. Deng, D. Xie, and W. Liu, “Cross-modal learning with adversarial samples,” vol. 32, 2019
2019
-
[47]
Adversarial attack on deep cross-modal hamming retrieval,
C. Li, S. Gao, C. Deng, W. Liu, and H. Huang, “Adversarial attack on deep cross-modal hamming retrieval,” in Int. Conf. Comput. Vis. , 2021, pp. 2218–2227
2021
-
[48]
Targeted adversarial attack against deep cross-modal hashing retrieval,
T. Wang, L. Zhu, Z. Zhang, H. Zhang, and J. Han, “Targeted adversarial attack against deep cross-modal hashing retrieval,” IEEE Trans. Circuit Syst. Video Technol., 2023
2023
-
[49]
Multifeature collaborative adversarial attack in multimodal remote sensing image classification,
C. Shi, Y . Dang, L. Fang, M. Zhao, Z. Lv, Q. Miao, and C.-M. Pun, “Multifeature collaborative adversarial attack in multimodal remote sensing image classification,” IEEE Trans. Inf. Geosci. Remote Sens. , vol. 60, pp. 1–15, 2022
2022
-
[50]
Unified adversarial patch for cross-modal attacks in the physical world,
X. Wei, Y . Huang, Y . Sun, and J. Yu, “Unified adversarial patch for cross-modal attacks in the physical world,” in Int. Conf. Comput. Vis. , 2023, pp. 4445–4454
2023
-
[51]
Enhancing adversarial example transferability with an intermediate level attack,
Q. Huang, I. Katsman, H. He, Z. Gu, S. Belongie, and S.-N. Lim, “Enhancing adversarial example transferability with an intermediate level attack,” in Int. Conf. Comput. Vis. , 2019, pp. 4733–4742
2019
-
[52]
Feature importance-aware transferable adversarial attacks,
Z. Wang, H. Guo, Z. Zhang, W. Liu, Z. Qin, and K. Ren, “Feature importance-aware transferable adversarial attacks,” in Int. Conf. Comput. Vis., 2021, pp. 7639–7648
2021
-
[53]
Enhancing the transferability of adversarial examples with random patch,
Y . Zhang, Y .-a. Tan, T. Chen, X. Liu, Q. Zhang, and Y . Li, “Enhancing the transferability of adversarial examples with random patch,” in Int. Jnt. Conf. Artif. Intell. , 2022, pp. 1672–1678
2022
-
[54]
Bag of tricks and a strong baseline for deep person re-identification,
H. Luo, Y . Gu, X. Liao, S. Lai, and W. Jiang, “Bag of tricks and a strong baseline for deep person re-identification,” in IEEE Conf. Comput. Vis. Pattern Recog. Worksh., 2019
2019
-
[55]
Unlabeled samples generated by gan improve the person re-identification baseline in vitro,
Z. Zheng, L. Zheng, and Y . Yang, “Unlabeled samples generated by gan improve the person re-identification baseline in vitro,” in Int. Conf. Comput. Vis., 2017, pp. 3754–3762
2017
-
[56]
Multi-scale deep learning architectures for person re-identification,
X. Qian, Y . Fu, Y .-G. Jiang, T. Xiang, and X. Xue, “Multi-scale deep learning architectures for person re-identification,” in Int. Conf. Comput. Vis., 2017, pp. 5399–5408
2017
-
[57]
Alignedreid: Surpassing human-level perfor- mance in person re-identification,
X. Zhang, H. Luo, X. Fan, W. Xiang, Y . Sun, Q. Xiao, W. Jiang, C. Zhang, and J. Sun, “Alignedreid: Surpassing human-level perfor- mance in person re-identification,” arXiv preprint arXiv:1711.08184 , 2017
2017 arXiv
-
[58]
Learning discriminative features with multiple granularities for person re-identification,
G. Wang, Y . Yuan, X. Chen, J. Li, and X. Zhou, “Learning discriminative features with multiple granularities for person re-identification,” in ACM Int. Conf. Multimedia , 2018, pp. 274–282
2018
-
[59]
Harmonious attention network for person re-identification,
W. Li, X. Zhu, and S. Gong, “Harmonious attention network for person re-identification,” in IEEE Conf. Comput. Vis. Pattern Recog., 2018, pp. 2285–2294
2018
-
[60]
Part-aware transformer for generalizable person re-identification,
H. Ni, Y . Li, L. Gao, H. T. Shen, and J. Song, “Part-aware transformer for generalizable person re-identification,” in Int. Conf. Comput. Vis. , 2023, pp. 11 280–11 289
2023
-
[61]
Channel augmentation for visible- infrared re-identification,
M. Ye, Z. Wu, C. Chen, and B. Du, “Channel augmentation for visible- infrared re-identification,” IEEE Trans. Pattern Anal. Mach. Intell. , vol. 46, no. 4, pp. 2299–2315, 2024
2024
-
[62]
Channel augmented joint learning for visible-infrared recognition,
M. Ye, W. Ruan, B. Du, and M. Z. Shou, “Channel augmented joint learning for visible-infrared recognition,” in Int. Conf. Comput. Vis. , 2021, pp. 13 567–13 576
2021
-
[63]
Dynamic dual-attentive aggregation learning for visible-infrared person re- identification,
M. Ye, J. Shen, D. J. Crandall, L. Shao, and J. Luo, “Dynamic dual-attentive aggregation learning for visible-infrared person re- identification,” in Eur. Conf. Comput. Vis., 2020, pp. 229–247
2020
-
[64]
Learning progressive modality-shared transformers for effective visible-infrared person re-identification,
H. Lu, X. Zou, and P. Zhang, “Learning progressive modality-shared transformers for effective visible-infrared person re-identification,” in AAAI Conf. Artif. Intell. , vol. 37, no. 2, 2023, pp. 1835–1843
2023
-
[65]
Towards a unified middle modality learning for visible-infrared person re-identification,
Y . Zhang, Y . Yan, Y . Lu, and H. Wang, “Towards a unified middle modality learning for visible-infrared person re-identification,” in ACM Int. Conf. Multimedia , 2021, pp. 788–796
2021
-
[66]
Unicat: Crafting a stronger fusion baseline for multimodal re-identification,
J. Crawford, H. Yin, L. McDermott, and D. Cummings, “Unicat: Crafting a stronger fusion baseline for multimodal re-identification,” arXiv preprint arXiv:2310.18812 , 2023
2023 arXiv
-
[67]
Deep residual learning for image recognition,
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in IEEE Conf. Comput. Vis. Pattern Recog., 2016, pp. 770– 778
2016
-
[68]
An image is worth 16x16 words: Transformers for image recognition at scale,
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly et al., “An image is worth 16x16 words: Transformers for image recognition at scale,” in Int. Conf. Learn. Represent. , 2020
2020
-
[69]
Densely connected convolutional networks,
G. Huang, Z. Liu, L. van der Maaten, and K. Q. Weinberger, “Densely connected convolutional networks,” in IEEE Conf. Comput. Vis. Pattern Recog., July 2017
2017
-
[70]
Rethinking the inception architecture for computer vision,
C. Szegedy, V . Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna, “Rethinking the inception architecture for computer vision,” in IEEE Conf. Comput. Vis. Pattern Recog., 2016, pp. 2818–2826
2016
-
[71]
Scalable person re-identification: A benchmark,
L. Zheng, L. Shen, L. Tian, S. Wang, J. Wang, and Q. Tian, “Scalable person re-identification: A benchmark,” in Int. Conf. Comput. Vis., 2015, pp. 1116–1124
2015
-
[72]
Diverse embedding expansion network and low-light cross-modality benchmark for visible-infrared person re- identification,
Y . Zhang and H. Wang, “Diverse embedding expansion network and low-light cross-modality benchmark for visible-infrared person re- identification,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2023, pp. 2153–2162
2023
-
[73]
Adam: a method for stochastic optimization,
D. Kingma, “Adam: a method for stochastic optimization,” in Int. Conf. Learn. Represent., 2014
2014
-
[74]
Performance measures and a data set for multi-target, multi-camera tracking,
E. Ristani, F. Solera, R. Zou, R. Cucchiara, and C. Tomasi, “Performance measures and a data set for multi-target, multi-camera tracking,” in Eur. Conf. Comput. Vis., 2016, pp. 17–35
2016
-
[75]
Keeping the bad guys out: Protecting and vaccinating deep learning with jpeg compression,
N. Das, M. Shanbhogue, S.-T. Chen, F. Hohman, L. Chen, M. E. Kounavis, and D. H. Chau, “Keeping the bad guys out: Protecting and vaccinating deep learning with jpeg compression,” arXiv preprint arXiv:1705.02900, 2017
2017 arXiv
-
[76]
Mitigating adversarial effects through randomization,
C. Xie, J. Wang, Z. Zhang, Z. Ren, and A. Yuille, “Mitigating adversarial effects through randomization,” in Int. Conf. Learn. Represent. , 2018
2018
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.