Pith. sign in

REVIEW 1 cited by

Magic Tokens: Select Diverse Tokens for Multi-modal Object Re-Identification

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.10254 v1 pith:UYW5YHLR submitted 2024-03-15 cs.CV cs.IRcs.MM

classification cs.CVcs.IRcs.MM
keywords multi-modalobjectreidtokensdiversefeaturemodalitiesselect
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Single-modal object re-identification (ReID) faces great challenges in maintaining robustness within complex visual scenarios. In contrast, multi-modal object ReID utilizes complementary information from diverse modalities, showing great potentials for practical applications. However, previous methods may be easily affected by irrelevant backgrounds and usually ignore the modality gaps. To address above issues, we propose a novel learning framework named \textbf{EDITOR} to select diverse tokens from vision Transformers for multi-modal object ReID. We begin with a shared vision Transformer to extract tokenized features from different input modalities. Then, we introduce a Spatial-Frequency Token Selection (SFTS) module to adaptively select object-centric tokens with both spatial and frequency information. Afterwards, we employ a Hierarchical Masked Aggregation (HMA) module to facilitate feature interactions within and across modalities. Finally, to further reduce the effect of backgrounds, we propose a Background Consistency Constraint (BCC) and an Object-Centric Feature Refinement (OCFR). They are formulated as two new loss functions, which improve the feature discrimination with background suppression. As a result, our framework can generate more discriminative features for multi-modal object ReID. Extensive experiments on three multi-modal ReID benchmarks verify the effectiveness of our methods. The code is available at https://github.com/924973292/EDITOR.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Modality Unified Attack for Omni-Modality Person Re-Identification

    cs.CV 2025-01 conditional novelty 6.0 of 10

    Modality-specific adversarial generators trained with metric disruption, simulated cross-modal, and collaborative multi-modal losses transfer to black-box single-, cross-, and multi-modality person re-id models, reach...

Pith tools