REVIEW 3 major objections 5 minor 79 references
Dual-Space Modality Consistency Learning for Universal Cross-Modal Re-Identification
T0 review · 3 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read Two training-only losses raise cross-modal ReID accuracy on all seventeen tested protocols.
desk verdict A well-run plug-and-play training method with consistent gains, but its core frequency-selective story is under-verified. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central machinery is a dual-space training objective rather than a new architecture. SMCL estimates, for each identity and modality in the mini-batch, a Gaussian with a mean vector and a diagonal covariance, then minimizes the resulting closed-form Wasserstein-2 distance, which reduces to the sum of squared Euclidean distances between the means and between the standard-deviation vectors. FDCL applies a 2D discrete Fourier transform to intermediate feature maps, computes the log-amplitude spectrum, shifts and standardizes it, and multiplies it by a lightweight learnable sigmoid mask that selects high-frequency regions; the selected vector is then trained with a supervised contrastive loss over cross-modal same-identity positives and different-identity negatives. The two losses are added, with a warm-up schedule, to the baseline's identity classification and triplet losses, and both are discarded at inference.
What would settle it
Run DSMCL on a version of SYSU-MM01 where all images are heavily Gaussian-blurred before feature extraction; if the Rank-1 gain over the baseline persists or grows, the paper's claim that FDCL works by preserving high-frequency identity cues would be falsified, because those cues no longer exist.
Extended reading notes
Core claim
The paper's central claim is that jointly optimizing spatial distribution consistency and frequency-domain discriminative consistency produces identity-consistent, modality-stable embeddings for cross-modal retrieval in a single framework that does not need modality-specific design. SMCL models each identity-modality class as a Gaussian, approximates the covariance as diagonal, and minimizes the simplified Wasserstein-2 distance between the mean vectors and standard deviation vectors of the same identity across modalities. FDCL transforms intermediate feature maps with a 2D discrete Fourier transform, takes the log-amplitude spectrum, standardizes it, extracts high-frequency regions with a learnable sigmoid mask, and runs supervised contrastive learning that pulls together same-identity samples from different modalities while pushing apart different identities. The two losses are trained alongside identity classification and triplet losses; at test time they are removed, so the backbone is unchanged. The paper validates the claim on seventeen evaluation protocols covering visible-infrared person ReID, optical-SAR ship ReID, and visible-NIR-TIR ship ReID, reporting improvements such as IDKL Rank-1 from 88.2% to 92.3% on SYSU-MM01 Indoor-search Single-shot and TransOSS from 31.3% to 38.8% on HOSS-ReID SAR-to-Optical.
Load-bearing premise
The argument depends on the assumption that the same high-frequency components that make identities distinguishable also carry most modality-specific distortion, so a contrastive loss on the amplitude spectrum can suppress one without destroying the other.
Editorial extensions
If this is right
- Adding DSMCL to an existing cross-modal ReID model should raise accuracy on visible-infrared, optical-SAR, and multi-spectral ship retrieval without changing the model used at inference time.
- Because the FDCL mask is learned, the framework should adapt to modalities with very different frequency statistics, such as infrared smoothing versus SAR speckle, rather than requiring hand-set frequency partitions.
- The gains reported on a frequency-aware baseline (MFENet) imply that frequency representation learning and explicit cross-modal frequency consistency are complementary, so stacking DSMCL on other frequency-aware models should also help.
- The class-wise Gaussian alignment of SMCL gives a general alternative to global distribution matching, so models that currently use adversarial or global alignment could swap in SMCL and expect more discriminative alignment.
Reading between the lines
- Editorial inference: the same dual-space recipe could transfer to other retrieval tasks with strong sensor gaps, such as text-image or sketch-photo person retrieval, since the losses are agnostic to the object category.
- Editorial inference: the paper's reported spectrum-discrepancy reductions suggest a diagnostic use: monitoring the difference between modality-averaged amplitude spectra could serve as a cheap early-stopping or domain-shift indicator in cross-modal systems.
- Editorial inference: a direct testable extension is to replace the supervised identity contrastive loss with self-supervised instance contrast on unlabeled cross-modal data; if the high-frequency mask carries genuine identity content, DSMCL-style pretraining should work without identity labels.
- Editorial inference: because FDCL operates on the amplitude spectrum only and leaves phase untouched, an implicit claim is that phase carries most of the spatial structure shared across modalities; if phase alignment matters more than the paper suggests, combining FDCL with lightweight phase regularization might yield further gains.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes DSMCL, a training-only, plug-and-play framework for cross-modal Re-Identification (ReID). It combines a spatial modality consistency loss (SMCL) that aligns per-class feature distributions across modalities using diagonal-Gaussian Wasserstein-2 distances, and a frequency-aware discriminative consistency loss (FDCL) that applies a 2D FFT to intermediate feature maps, learns a soft frequency mask, and performs identity-aware contrastive learning on the masked log-amplitude spectrum. The total loss is a weighted sum of identity, triplet, SMCL, and FDCL terms. The method is evaluated on SYSU-MM01, RegDB, LLCM, HOSS-ReID, and CMShipReID, using IDKL, MFENet, HSFLNet, TransOSS, MOS, and TransReID as baselines, and reports consistent improvements across seventeen evaluation protocols.
Significance. If the empirical claims hold, the contribution is practically useful: DSMCL improves multiple representative baselines across person and ship cross-modal ReID without changing inference cost. The loss formulations in Eqs. (10), (21), and (24) are mathematically coherent, and the evaluation breadth is substantial. The paper does not provide code, and all reported numbers are single-run point estimates without error bars or statistical tests, which limits confidence in the magnitude and consistency of the claimed gains. The central explanatory mechanism, that FDCL selectively removes modality-specific high-frequency distortions while preserving identity-bearing high-frequency cues, is not directly verified.
major comments (3)
- [Section III.C.2, Eq. (18); Section IV.F.2] The paper's core claim is that FDCL suppresses modality-specific high-frequency distortions while preserving identity-discriminative high-frequency details. This relies on the assumption stated in Section III.C that high-frequency components contain both discriminative identity information and modality-sensitive distortions. However, the mask in Eq. (18), M_high = Sigmoid(G(A_hat_i)), has no constraint, regularization, or architectural bias that forces it to select high-frequency regions; it can select any spectral region that lowers the contrastive loss, including low-frequency or modality-discriminative artifacts. The only supporting evidence, Section IV.F.2, shows that after training the cross-modal amplitude discrepancy is reduced; this is an output effect of any contrastive alignment and does not demonstrate that the retained coefficients are identity-bearing. Please provide direct evidence: (a) report the learned mask's frequency support, for example the average mask energy as a function of radial frequency; (b) compare FDCL against a fixed high-frequency mask and against a full-spectrum contrastive baseline; (c) evaluate the identity discriminability of the masked spectrum, for example by linear classification or retrieval on the masked representation. Without such evidence, the reported gains could come from generic contrastive regularization on the amplitude spectrum, and the frequency-aware explanation, which is the paper's main novelty, remains unsupported.
- [Tables I-V; Section IV.B] All experimental results are single-run point estimates without error bars, multiple seeds, or paired significance tests. Many reported gains are modest, for example IDKL on SYSU-MM01 All-search Single-shot improves from 80.9% to 83.6% Rank-1, and several CMShipReID cells change by less than two percentage points. The claim that DSMCL 'consistently improves' all baselines across seventeen protocols therefore cannot be assessed statistically. In addition, the loss weights lambda_s and lambda_f are selected on validation performance and the temperature tau is fixed at 0.2 based on RegDB (Section IV.B, Section IV.E); please state explicitly whether the same hyperparameter values are used for every dataset and protocol, and report the variance of the results over at least three independent runs with paired comparison against each baseline.
- [Section III.C.3, Eq. (21)] The FDCL contrastive loss in Eq. (21) assumes that positive sets P(i) are non-empty, but in mini-batch sampling an identity may appear in only one modality, especially on the smaller HOSS-ReID and CMShipReID datasets. The paper defines V as the 'valid anchor set' but does not specify how anchors with empty P(i) are handled or whether such samples are excluded from the loss. Please clarify the exact construction of P(i), N(i), and V, including whether the denominator excludes the anchor itself and whether same-modality same-identity samples are intentionally excluded from positives. This affects the practical implementation and the comparability of the loss across datasets with different modality counts and batch compositions.
minor comments (5)
- [Section III.B, Eq. (10)] The sum over (m_a, m_b) in Eq. (10) is not fully specified; for datasets with three modalities, please state explicitly whether the sum runs over ordered or unordered distinct modality pairs and how terms with m_a = m_b are treated.
- [Section IV.B] The progressive warm-up schedule is described as activating SMCL after 20 epochs and FDCL after 40 epochs with a linear ramp over 40 epochs, but the total number of training epochs is not stated anywhere in the paper. Please report the epoch counts and the exact ramp schedule used for each baseline.
- [Section IV.F.2] The frequency spectrum alignment analysis uses only ten randomly selected identities and reports no variability across selections or runs. Please state how the identities were selected, whether the result is stable across different selections, and consider reporting the analysis over the full test set.
- [Figure 10] The caption contains a duplicated word: 'between between' should be corrected to 'between'.
- [Section III.C.1] The logarithmic amplitude spectrum in Eq. (15) is standardized per image in Eq. (17), but it is not discussed how the standardization interacts with the subsequent mask and contrastive similarity; a one-sentence justification or an ablation on this normalization would improve clarity.
Circularity Check
No significant circularity: DSMCL's claimed gains are external benchmark results, not derivations from its own inputs.
full rationale
The paper's central claim is empirical: adding SMCL and FDCL losses to existing cross-modal ReID models improves Rank-1/mAP on SYSU-MM01, RegDB, LLCM, HOSS-ReID, and CMShipReID across 17 protocols. These results are measured on held-out test splits against independently reproduced baselines (IDKL, MFENet, HSFLNet, TransOSS, MOS, TransReID), so the reported improvements are not forced by construction from the training objectives. The frequency-domain premise in Sec. III.C ('heterogeneous imaging mechanisms predominantly affect high-frequency responses while discriminative identity information is also mainly concentrated in high-frequency components') is a stated assumption, not a conclusion derived from the loss; the learnable mask in Eq. (18) is indeed unconstrained, and the post-hoc spectrum analysis in Sec. IV.F.2 is an output effect rather than proof that identity cues survive, but these are correctness/validation gaps, not circularity. Hyperparameters (lambda_s, lambda_f, tau) are selected on validation data and kept fixed for test evaluation, which is standard practice; no fitted parameter is renamed as a prediction. The paper cites the authors' prior work (MOS, BIT, CUUP, Try Harder) and explicitly positions DSMCL as an extension of MOS, but these citations are not load-bearing for the benchmark comparisons: MOS is treated as an external baseline with reproduced numbers, and DSMCL's improvement over MOS is a measured outcome, not a consequence of citing MOS. No uniqueness theorem or self-citation chain forces the DSMCL design. Hence, no definitional, fitted-input, or self-citation circularity is present; the paper is self-contained against external benchmarks.
Assumptions & free parameters
free parameters (4)
- λ_s (spatial consistency loss weight) =
0.05
- λ_f (frequency consistency loss weight) =
0.03
- τ (FDCL temperature) =
0.2
- Progressive warm-up schedule =
SMCL after 20 epochs; FDCL after 40 epochs with 40-epoch linear ramp
assumptions (5)
- domain assumption Class-conditional feature distributions can be approximated by multivariate Gaussians via the central limit theorem.
- ad hoc to paper Full covariance matrices can be replaced by diagonal variances without losing distributional expressiveness needed for alignment.
- domain assumption Modality discrepancy is primarily reflected in the amplitude spectrum, not the phase, and high-frequency components carry both discriminative identity cues and modality-sensitive distortions.
- ad hoc to paper The learnable mask generator learns to select informative high-frequency regions rather than trivial or noisy regions.
- domain assumption Baseline implementations and reproduced numbers are fair and correctly configured.
invented entities (1)
-
Learnable frequency mask generator G(·)
Cite this review
Pith. "Pith review of Dual-Space Modality Consistency Learning for Universal Cross-Modal Re-Identification." pith.science (2026). https://pith.science/paper/4BV3YNTR
@misc{pith2026260806943,
author = {Pith},
title = {Pith review of: Dual-Space Modality Consistency Learning for Universal Cross-Modal Re-Identification},
year = {2026},
howpublished = {\url{https://pith.science/paper/4BV3YNTR}},
note = {Machine review of arXiv:2608.06943}
}
read the original abstract
Cross-modal Re-Identification (ReID) aims to retrieve the same identity across heterogeneous imaging modalities and has been widely studied in visible-infrared person ReID and cross-modal ship ReID. Existing methods have achieved promising performance by learning modality consistency in the spatial embedding space, yet often overlook frequency-domain modality discrepancy, particularly in high-frequency representations that are both highly discriminative and modality-sensitive. In addition, most approaches are tailored to specific modality settings, limiting their applicability across diverse cross-modal scenarios. To address these challenges, we propose a Dual-Space Modality Consistency Learning (DSMCL) framework for universal cross-modal ReID. Specifically, DSMCL jointly models spatial feature distribution consistency and frequency-domain discriminative consistency. A Spatial Modality Consistency Learning (SMCL) branch performs Gaussian-based feature alignment, while a Frequency-aware Discriminative Consistency Learning (FDCL) strategy regularizes high-frequency representations through identity-aware cross-modal contrastive learning. By jointly capturing modality-specific characteristics and modality-shared identity cues, DSMCL learns robust representations and establishes a unified framework capable of accommodating diverse heterogeneous modality settings. Moreover, DSMCL is a plug-and-play framework that can be readily integrated into existing cross-modal ReID architectures. Extensive experiments on SYSU-MM01, RegDB, LLCM, HOSS-ReID, and CMShipReID across seventeen evaluation protocols show that DSMCL consistently improves multiple representative baselines.
Figures
Figures from the paper (6 more)
Reference graph
Works this paper leans on
-
[1]
S. Li, L. Sun, and Q. Li, “Clip-reid: exploiting vision-language model for image re-identification without concrete text labels,” inProc. Conf. Assoc. Adv. Artif. Intell., vol. 37, no. 1, 2023, pp. 1405–1413
work page 2023
-
[2]
Infrared-visible cross-modal person re-identification with an x modality,
D. Li, X. Wei, X. Hong, and Y . Gong, “Infrared-visible cross-modal person re-identification with an x modality,” inProc. Conf. Assoc. Adv. Artif. Intell., vol. 34, no. 4, 2020, pp. 4610–4617
work page 2020
-
[3]
Instruct-reid: A multi-purpose person re-identification task with instructions,
W. Heet al., “Instruct-reid: A multi-purpose person re-identification task with instructions,” inProc. IEEE Conf. Comput. Vis. Pattern Recognit., 2024, pp. 17 521–17 531
work page 2024
-
[4]
Cloth-changing person re-identification from a single image with gait prediction and regularization,
X. Jinet al., “Cloth-changing person re-identification from a single image with gait prediction and regularization,” inProc. IEEE Conf. Comput. Vis. Pattern Recognit., 2022, pp. 14 278–14 287
work page 2022
-
[5]
Try harder: Hard sample generation and learning for cloth-changing person re-id,
H. Liu, Y . Zhao, and G. Niu, “Try harder: Hard sample generation and learning for cloth-changing person re-id,” inProc. ACM Int. Conf. Multimedia, 2025, pp. 1704–1713
work page 2025
-
[6]
Y . Zhao, C. Wu, Y . Xu, X. Du, R. Li, and G. Niu, “Ccup: A controllable synthetic data generation pipeline for pretraining cloth-changing person re-identification models,” inProc. IEEE Int. Conf. on Multimedia and Expo, 2025, pp. 1–6
work page 2025
-
[7]
Towards Robust Text-to-Image Person Retrieval: Multi-View Reformulation for Semantic Compensation
C. Yuan, Y . Zhao, H. Xu, and G. Niu, “Towards robust text-to-image person retrieval: Multi-view reformulation for semantic compensation,” arXiv:2604.18376, 2026
work page Pith review arXiv 2026
-
[8]
N. Huang, J. Liu, Y . Miao, Q. Zhang, and J. Han, “Deep learning for visible-infrared cross-modality person re-identification: A comprehen- sive review,”Information Fusion, vol. 91, pp. 396–411, 2023
work page 2023
Show all 79 references
-
[9]
Diverse embedding expansion network and low-light cross-modality benchmark for visible-infrared person re- identification,
Y . Zhang and H. Wang, “Diverse embedding expansion network and low-light cross-modality benchmark for visible-infrared person re- identification,” inProc. IEEE Conf. Comput. Vis. Pattern Recognit., 2023, pp. 2153–2162
2023
-
[10]
X-reid: Multi-granularity informa- tion interaction for video-based visible-infrared person re-identification,
C. Yu, X. Liu, P. Zhang, and H. Lu, “X-reid: Multi-granularity informa- tion interaction for video-based visible-infrared person re-identification,” inProc. Conf. Assoc. Adv. Artif. Intell., vol. 40, no. 14, 2026, pp. 12 117– 12 125
2026
-
[11]
Shape-erased feature learning for visible-infrared person re-identification,
J. Feng, A. Wu, and W.-S. Zheng, “Shape-erased feature learning for visible-infrared person re-identification,” inProc. IEEE Conf. Comput. Vis. Pattern Recognit., 2023, pp. 22 752–22 761
2023
-
[12]
Cross-modality person re- identification via modality confusion and center aggregation,
X. Hao, S. Zhao, M. Ye, and J. Shen, “Cross-modality person re- identification via modality confusion and center aggregation,” inProc. IEEE Int. Conf. Comput. Vis., 2021, pp. 16 403–16 412
2021
-
[13]
Neural feature search for rgb-infrared person re-identification,
Y . Chen, L. Wan, Z. Li, Q. Jing, and Z. Sun, “Neural feature search for rgb-infrared person re-identification,” inProc. IEEE Conf. Comput. Vis. Pattern Recognit., 2021, pp. 587–597
2021
-
[14]
Cross-modal ship re-identification via optical and sar imagery: A novel dataset and method,
H. Wang, S. Li, J. Yang, Y . Liu, Y . Lv, and Z. Zhou, “Cross-modal ship re-identification via optical and sar imagery: A novel dataset and method,” inProc. IEEE Int. Conf. Comput. Vis., 2025, pp. 7873–7883
2025
-
[15]
Mos: Mitigating optical-sar modality gap for cross-modal ship re-identification,
Y . Zhao, H. Liu, and G. Niu, “Mos: Mitigating optical-sar modality gap for cross-modal ship re-identification,” inProc. IEEE Conf. Comput. Vis. Pattern Recognit., 2026, pp. 30 335–30 345
2026
-
[16]
Smart-ship: A comprehensive synchronized multi- modal aligned remote sensing targets dataset and benchmark for berthed ships analysis,
C.-C. Fanet al., “Smart-ship: A comprehensive synchronized multi- modal aligned remote sensing targets dataset and benchmark for berthed ships analysis,”arXiv:2508.02384, 2025
2025 arXiv
-
[17]
Cmshipreid: A cross-modality ship dataset for the re- identification task,
C. Xuet al., “Cmshipreid: A cross-modality ship dataset for the re- identification task,”IEEE J. Sel. Top. Appl. Earth Observ. Remote Sens., vol. 18, pp. 10 503–10 513, 2025
2025
-
[18]
Modality-transition representation learning for visible- infrared person re-identification,
C. Yuanet al., “Modality-transition representation learning for visible- infrared person re-identification,”arXiv:2511.02685, 2025
2025
-
[19]
Diffusion-based synthetic data generation for visible-infrared person re-identification,
W. Dai, L. Lu, and Z. Li, “Diffusion-based synthetic data generation for visible-infrared person re-identification,” inProc. Conf. Assoc. Adv. Artif. Intell., vol. 39, no. 11, 2025, pp. 11 185–11 193
2025
-
[20]
A generative-based image fusion strategy for visible-infrared person re-identification,
J. Qi, T. Liang, W. Liu, Y . Li, and Y . Jin, “A generative-based image fusion strategy for visible-infrared person re-identification,”IEEE Trans. Circuits Syst. Video Technol., vol. 34, no. 1, pp. 518–533, 2023
2023
-
[21]
Unified conditional image generation for visible-infrared person re-identification,
H. Pan, W. Pei, X. Li, and Z. He, “Unified conditional image generation for visible-infrared person re-identification,”IEEE Trans. Inf. F orensics Security, vol. 19, pp. 9026–9038, 2024
2024
-
[22]
Learning modality-specific representa- tions for visible-infrared person re-identification,
Z. Feng, J. Lai, and X. Xie, “Learning modality-specific representa- tions for visible-infrared person re-identification,”IEEE Trans. Image Process., vol. 29, pp. 579–590, 2019
2019
-
[23]
Identity-compensated style distillation for visible-infrared person re-identification,
Y . Ling, Z. Hu, N. Pu, Z. Zhong, and X. Jiang, “Identity-compensated style distillation for visible-infrared person re-identification,”IEEE Trans. Image Process., vol. 35, pp. 2941–2954, 2026
2026
-
[24]
Learning progressive modality-shared transformers for effective visible-infrared person re-identification,
H. Lu, X. Zou, and P. Zhang, “Learning progressive modality-shared transformers for effective visible-infrared person re-identification,” in Proc. Conf. Assoc. Adv. Artif. Intell., vol. 37, no. 2, 2023, pp. 1835– 1843
2023
-
[25]
Visible-infrared person re-identification via cross- modality interaction transformer,
Y . Fenget al., “Visible-infrared person re-identification via cross- modality interaction transformer,”IEEE Trans. Multimedia, vol. 25, pp. 7647–7659, 2022
2022
-
[26]
Dma: Dual modality-aware alignment for visible-infrared person re-identification,
Z. Cui, J. Zhou, and Y . Peng, “Dma: Dual modality-aware alignment for visible-infrared person re-identification,”IEEE Trans. Inf. F orensics Security, vol. 19, pp. 2696–2708, 2024
2024
-
[27]
Adaptive middle modality alignment learning for visible-infrared person re-identification,
Y . Zhang, Y . Yan, Y . Lu, and H. Wang, “Adaptive middle modality alignment learning for visible-infrared person re-identification,”Int. J. Comput. Vis., vol. 133, no. 4, pp. 2176–2196, 2025
2025
-
[28]
Diverse semantics- guided feature alignment and decoupling for visible-infrared person re- identification,
N. Dong, S. Yan, L. Zhang, and J. Tang, “Diverse semantics- guided feature alignment and decoupling for visible-infrared person re- identification,”IEEE Trans. Inf. F orensics Security, vol. 20, pp. 12 245– 12 259, 2025
2025
-
[29]
Dsaf: Dual space alignment framework for visible-infrared person re-identification,
Y . Jiang, X. Cheng, H. Yu, X. Liu, H. Chen, and G. Zhao, “Dsaf: Dual space alignment framework for visible-infrared person re-identification,” IEEE Trans. Multimedia, vol. 27, pp. 5591–5603, 2025
2025
-
[30]
Propagation based recycling contrastive learning for coupled noisy visible-infrared person re-identification,
Y . Li, W. Tang, S. Wang, S. Qian, Q. Fang, and C. Xu, “Propagation based recycling contrastive learning for coupled noisy visible-infrared person re-identification,”IEEE Trans. Multimedia, vol. 27, pp. 9330– 9341, 2025
2025
-
[31]
Not all pixels are matched: Dense contrastive learning for cross-modality person re-identification,
H. Sunet al., “Not all pixels are matched: Dense contrastive learning for cross-modality person re-identification,” inProc. ACM Int. Conf. Multimedia, 2022, pp. 5333–5341
2022
-
[32]
Attend to the difference: Cross-modality person re-identification via contrastive correlation,
S. Zhang, Y . Yang, P. Wang, G. Liang, X. Zhang, and Y . Zhang, “Attend to the difference: Cross-modality person re-identification via contrastive correlation,”IEEE Trans. Image Process., vol. 30, pp. 8861–8872, 2021
2021
-
[33]
Clip-driven semantic discovery network for visible-infrared person re-identification,
X. Yu, N. Dong, L. Zhu, H. Peng, and D. Tao, “Clip-driven semantic discovery network for visible-infrared person re-identification,”IEEE Trans. Multimedia, vol. 27, pp. 4137–4150, 2025
2025
-
[34]
Bridging the gap: Multi-level cross-modality joint alignment for visible-infrared person re- identification,
T. Liang, Y . Jin, W. Liu, T. Wang, S. Feng, and Y . Li, “Bridging the gap: Multi-level cross-modality joint alignment for visible-infrared person re- identification,”IEEE Trans. Circuits Syst. Video Technol., vol. 34, no. 8, pp. 7683–7698, 2024
2024
-
[35]
Prototype-driven hierarchical skeleton modeling for visible-infrared person re-identification,
J. Pan, B. Zhang, and X. Cao, “Prototype-driven hierarchical skeleton modeling for visible-infrared person re-identification,”Eng. Appl. Artif. Intell., vol. 164, p. 113224, 2026
2026
-
[36]
A general spatial-frequency learning framework for multimodal image fusion,
M. Zhouet al., “A general spatial-frequency learning framework for multimodal image fusion,”IEEE Trans. Pattern Anal. Mach. Intell., vol. 47, no. 7, pp. 5281–5298, 2024. JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021 17
2024
-
[37]
Frequency and spatial-domain saliency network for remote sensing cross-modal retrieval,
C. Zheng, J. Nie, B. Yin, X. Li, Y . Qian, and Z. Wei, “Frequency and spatial-domain saliency network for remote sensing cross-modal retrieval,”IEEE Trans. Geosci. Remote Sens., vol. 63, pp. 1–13, 2025
2025
-
[38]
Freqcross: A multi-modal frequency-spatial fusion network for robust detection of stable diffusion 3.5 generated images,
G. Yang, “Freqcross: A multi-modal frequency-spatial fusion network for robust detection of stable diffusion 3.5 generated images,” inProc. IEEE Conf. Comput. Vis. Pattern Recognit., 2025, pp. 6577–6584
2025
-
[39]
Adaptive high-frequency transformer for diverse wildlife re-identification,
C. Li, S. Chen, and M. Ye, “Adaptive high-frequency transformer for diverse wildlife re-identification,” inProc. Eur . Conf. Comput. Vis., 2024, pp. 296–313
2024
-
[40]
Discovering multi- frequency embedding for visible-infrared person re-identification,
H. Gu, X. Yang, R. Lu, L. Pu, S. Han, and M. Wu, “Discovering multi- frequency embedding for visible-infrared person re-identification,”IEEE Trans. Circuits Syst. Video Technol., vol. 36, no. 2, pp. 1766–1780, 2025
2025
-
[41]
Rgb-infrared cross- modality person re-identification,
A. Wu, W.-S. Zheng, H.-X. Yu, S. Gong, and J. Lai, “Rgb-infrared cross- modality person re-identification,” inProc. IEEE Int. Conf. Comput. Vis., 2017, pp. 5390–5399
2017
-
[42]
Person recognition system based on a combination of body images from visible light and thermal cameras,
D. T. Nguyen, H. G. Hong, K. W. Kim, and K. R. Park, “Person recognition system based on a combination of body images from visible light and thermal cameras,”Sensors, vol. 17, no. 3, p. 605, 2017
2017
-
[43]
Visible-infrared person re-identification via semantic alignment and affinity inference,
X. Fang, Y . Yang, and Y . Fu, “Visible-infrared person re-identification via semantic alignment and affinity inference,” inProc. IEEE Conf. Comput. Vis. Pattern Recognit., 2023, pp. 11 270–11 279
2023
-
[44]
A versatile framework for multi- scene person re-identification,
W.-S. Zheng, J. Yan, and Y .-X. Peng, “A versatile framework for multi- scene person re-identification,”IEEE Trans. Pattern Anal. Mach. Intell., vol. 47, no. 3, pp. 1362–1380, 2024
2024
-
[45]
Deep learning for person re-identification: A survey and outlook,
M. Ye, J. Shen, G. Lin, T. Xiang, L. Shao, and S. C. Hoi, “Deep learning for person re-identification: A survey and outlook,”IEEE Trans. Pattern Anal. Mach. Intell., vol. 44, no. 6, pp. 2872–2893, 2021
2021
-
[46]
Bit: Matching-based bi-directional interaction transformation network for visible-infrared person re-identification,
H. Xu and G. Niu, “Bit: Matching-based bi-directional interaction transformation network for visible-infrared person re-identification,” in Proc. IEEE Conf. Comput. Vis. Pattern Recognit., 2026, pp. 40 386– 40 396
2026
-
[47]
Hoh- net: High-order hierarchical middle-feature learning network for visible- infrared person re-identification,
L. Qiu, S. Chen, J.-H. Xue, D.-H. Wang, S. Zhu, and Y . Yan, “Hoh- net: High-order hierarchical middle-feature learning network for visible- infrared person re-identification,”IEEE Trans. Circuits Syst. Video Technol., vol. 36, no. 2, pp. 2607–2622, 2025
2025
-
[48]
Cross-modality self-distillation for visible-infrared person re-identification,
Z. Hu, Y . Ling, Z. Luo, W. Shao, S. Li, and Z. Zhong, “Cross-modality self-distillation for visible-infrared person re-identification,”IEEE Trans. Geosci. Remote Sens., 2026
2026
-
[49]
Adaptive generation of privileged intermediate information for visible- infrared person re-identification,
M. Alehdaghi, A. Josi, R. M. Cruz, P. Shamsolmoali, and E. Granger, “Adaptive generation of privileged intermediate information for visible- infrared person re-identification,”IEEE Trans. Inf. F orensics Security, vol. 20, pp. 3400–3413, 2025
2025
-
[50]
Beyond weight adaptation: Feature-space domain injection for cross-modal ship re-identification,
T. Xian, W. Zhou, Z. Zhou, and Z. Li, “Beyond weight adaptation: Feature-space domain injection for cross-modal ship re-identification,” arXiv:2512.20892, 2025
2025
-
[51]
Sdf-net: Structure-aware disentangled feature learning for opticall-sar ship re-identification,
F. Chenet al., “Sdf-net: Structure-aware disentangled feature learning for opticall-sar ship re-identification,”arXiv:2603.12588, 2026
2026 arXiv
-
[52]
Learning in the frequency domain,
K. Xu, M. Qin, F. Sun, Y . Wang, Y .-K. Chen, and F. Ren, “Learning in the frequency domain,” inProc. IEEE Conf. Comput. Vis. Pattern Recognit., 2020, pp. 1740–1749
2020
-
[53]
Thinking in frequency: Face forgery detection by mining frequency-aware clues,
Y . Qian, G. Yin, L. Sheng, Z. Chen, and J. Shao, “Thinking in frequency: Face forgery detection by mining frequency-aware clues,” inProc. Eur . Conf. Comput. Vis., 2020, pp. 86–103
2020
-
[54]
Pha: Patch-wise high- frequency augmentation for transformer-based person re-identification,
G. Zhang, Y . Zhang, T. Zhang, B. Li, and S. Pu, “Pha: Patch-wise high- frequency augmentation for transformer-based person re-identification,” inProc. IEEE Conf. Comput. Vis. Pattern Recognit., 2023, pp. 14 133– 14 142
2023
-
[55]
Mfen: Multi-frequency expert network for visible-infrared person re-id,
X. Liet al., “Mfen: Multi-frequency expert network for visible-infrared person re-id,” inProc. IEEE Conf. Comput. Vis. Pattern Recognit., 2026, pp. 18 471–18 480
2026
-
[56]
Frequency-driven feature decoupling network for visible–infrared person re-identification,
X. Dong, Q. Tan, R. Wang, X. Liu, and P. Cao, “Frequency-driven feature decoupling network for visible–infrared person re-identification,” Eng. Appl. Artif. Intell., vol. 167, p. 113740, 2026
2026
-
[57]
Frequency domain nu- ances mining for visible-infrared person re-identification,
Y . Zhang, H. Wang, Y . Lu, Y . Yan, and X. Li, “Frequency domain nu- ances mining for visible-infrared person re-identification,”IEEE Trans. Inf. F orensics Security, vol. 20, pp. 5411–5424, 2025
2025
-
[58]
Partmix: Regular- ization strategy to learn part discovery for visible-infrared person re- identification,
M. Kim, S. Kim, J. Park, S. Park, and K. Sohn, “Partmix: Regular- ization strategy to learn part discovery for visible-infrared person re- identification,” inProc. IEEE Conf. Comput. Vis. Pattern Recognit., 2023, pp. 18 621–18 632
2023
-
[59]
Modality unifying network for visible-infrared person re-identification,
H. Yu, X. Cheng, W. Peng, W. Liu, and G. Zhao, “Modality unifying network for visible-infrared person re-identification,” inProc. IEEE Int. Conf. Comput. Vis., 2023, pp. 11 185–11 195
2023
-
[60]
Channel augmentation for visible- infrared re-identification,
M. Ye, Z. Wu, C. Chen, and B. Du, “Channel augmentation for visible- infrared re-identification,”IEEE Trans. Pattern Anal. Mach. Intell., vol. 46, no. 4, pp. 2299–2315, 2023
2023
-
[61]
Variational distillation for multi-view learning,
X. Tianet al., “Variational distillation for multi-view learning,”IEEE Trans. Pattern Anal. Mach. Intell., vol. 46, no. 7, pp. 4551–4566, 2023
2023
-
[62]
High-order structure based middle-feature learning for visible-infrared person re- identification,
L. Qiu, S. Chen, Y . Yan, J. Xue, D. Wang, and S. Zhu, “High-order structure based middle-feature learning for visible-infrared person re- identification,” inProc. Conf. Assoc. Adv. Artif. Intell., vol. 38, no. 5, 2024, pp. 4596–4604
2024
-
[63]
Wrim-net: Wide-ranging information mining network for visible-infrared person re-identification,
Y . Wu, L.-C. Meng, Y . Zichao, S. Chan, and H.-Q. Wang, “Wrim-net: Wide-ranging information mining network for visible-infrared person re-identification,” inProc. Eur . Conf. Comput. Vis., 2024, pp. 55–72
2024
-
[64]
Domain shifting: A generalized solution for heterogeneous cross-modality person re-identification,
Y . Jiang, X. Cheng, H. Yu, X. Liu, H. Chen, and G. Zhao, “Domain shifting: A generalized solution for heterogeneous cross-modality person re-identification,” inProc. Eur . Conf. Comput. Vis., 2024, pp. 289–306
2024
-
[65]
Rle: A unified perspective of data augmentation for cross- spectral re-identification,
T. Leiet al., “Rle: A unified perspective of data augmentation for cross- spectral re-identification,”Proc. Adv. Conf. Neural Inf. Process. Syst., vol. 37, pp. 126 977–126 996, 2024
2024
-
[66]
Causality-invariant interactive mining for cross-modal similarity learning,
J. Yan, C. Deng, H. Huang, and W. Liu, “Causality-invariant interactive mining for cross-modal similarity learning,”IEEE Trans. Pattern Anal. Mach. Intell., vol. 46, no. 9, pp. 6216–6230, 2024
2024
-
[67]
Cooperative separation of modality shared-specific features for visible-infrared per- son re-identification,
X. Yang, W. Dong, M. Li, Z. Wei, N. Wang, and X. Gao, “Cooperative separation of modality shared-specific features for visible-infrared per- son re-identification,”IEEE Trans. Multimedia, vol. 26, pp. 8172–8183, 2024
2024
-
[68]
Advancing visible-infrared person re-identification: Synergizing visual-textual rea- soning and cross-modal feature alignment,
Y . Qiu, L. Wang, W. Song, J. Liu, Z. Shi, and N. Jiang, “Advancing visible-infrared person re-identification: Synergizing visual-textual rea- soning and cross-modal feature alignment,”IEEE Trans. Inf. F orensics Security, vol. 20, pp. 2184–2196, 2025
2025
-
[69]
Hierarchical token-aware cross-modality reconstruction for visible-infrared person re- identification,
S. Chen, L. Qiu, D.-H. Wang, W. Zhu, Y . Hua, and Y . Yan, “Hierarchical token-aware cross-modality reconstruction for visible-infrared person re- identification,”IEEE Trans. Multimedia, vol. 27, pp. 7012–7027, 2025
2025
-
[70]
Disentangling modality and posture factors: Memory-attention and orthogonal decomposition for visible-infrared person re-identification,
Z. Lu, R. Lin, and H. Hu, “Disentangling modality and posture factors: Memory-attention and orthogonal decomposition for visible-infrared person re-identification,”IEEE Trans. Neural Netw. Learn. Syst., vol. 36, no. 3, pp. 5494–5508, 2025
2025
-
[71]
Cycletrans: Learning neutral yet discriminative features via cycle construction for visible- infrared person re-identification,
Q. Wu, J. Xia, P. Dai, Y . Zhou, Y . Wu, and R. Ji, “Cycletrans: Learning neutral yet discriminative features via cycle construction for visible- infrared person re-identification,”IEEE Trans. Neural Netw. Learn. Syst., vol. 36, no. 3, pp. 5469–5479, 2025
2025
-
[72]
Modality-perceptive harmo- nization network for visible-infrared person re-identification,
X. Zuo, J. Peng, T. Cheng, and H. Wang, “Modality-perceptive harmo- nization network for visible-infrared person re-identification,”Informa- tion Fusion, vol. 118, p. 102979, 2025
2025
-
[73]
Uncertainty- aware cross-modal opinion interaction: A general framework for visible- infrared vehicle and person re-identification,
S. Shan, H. Liu, F. Shang, Q. Wang, and Y . Song, “Uncertainty- aware cross-modal opinion interaction: A general framework for visible- infrared vehicle and person re-identification,” inProc. IEEE Conf. Comput. Vis. Pattern Recognit., 2026, pp. 6476–6485
2026
-
[74]
Implicit discriminative knowledge learning for visible-infrared person re-identification,
K. Ren and L. Zhang, “Implicit discriminative knowledge learning for visible-infrared person re-identification,” inProc. IEEE Conf. Comput. Vis. Pattern Recognit., 2024, pp. 393–402
2024
-
[75]
Hypergraph-driven soft semantics flexible learning for visible–infrared person re-identification,
J. Zhu, H. Ge, Y . Liu, C. Wu, and J. Fan, “Hypergraph-driven soft semantics flexible learning for visible–infrared person re-identification,” Eng. Appl. Artif. Intell., vol. 158, p. 111286, 2025
2025
-
[76]
Mdanet: Modality-aware domain alignment network for visible-infrared person re-identification,
X. Cheng, H. Yu, K. H. M. Cheng, Z. Yu, and G. Zhao, “Mdanet: Modality-aware domain alignment network for visible-infrared person re-identification,”IEEE Trans. Multimedia, vol. 27, pp. 2015–2027, 2024
2015
-
[77]
Transreid: Transformer-based object re-identification,
S. He, H. Luo, P. Wang, F. Wang, H. Li, and W. Jiang, “Transreid: Transformer-based object re-identification,” inProc. IEEE Int. Conf. Comput. Vis., 2021, pp. 15 013–15 022
2021
-
[78]
Beyond appearance: a semantic controllable self- supervised learning framework for human-centric visual tasks,
W. Chenet al., “Beyond appearance: a semantic controllable self- supervised learning framework for human-centric visual tasks,” inProc. IEEE Conf. Comput. Vis. Pattern Recognit., 2023, pp. 15 050–15 061
2023
-
[79]
Advancing ship re-identification in the wild: The shipreid- 2400 benchmark dataset and d2internet baseline method,
B. Liuet al., “Advancing ship re-identification in the wild: The shipreid- 2400 benchmark dataset and d2internet baseline method,” inProc. Int. ACM SIGIR Conf. Res. Develop. Inf. Retrieval, 2025, pp. 106–115
2025
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.