Pith. sign in

REVIEW 5 cited by

GraFT: Gradual Fusion Transformer for Multimodal Re-Identification

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.16856 v1 pith:OCLMTGQG submitted 2023-10-25 cs.CV

classification cs.CV
keywords graftfusionmultimodalreidgradualintroducere-identificationtransformer
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Object Re-Identification (ReID) is pivotal in computer vision, witnessing an escalating demand for adept multimodal representation learning. Current models, although promising, reveal scalability limitations with increasing modalities as they rely heavily on late fusion, which postpones the integration of specific modality insights. Addressing this, we introduce the \textbf{Gradual Fusion Transformer (GraFT)} for multimodal ReID. At its core, GraFT employs learnable fusion tokens that guide self-attention across encoders, adeptly capturing both modality-specific and object-specific features. Further bolstering its efficacy, we introduce a novel training paradigm combined with an augmented triplet loss, optimizing the ReID feature embedding space. We demonstrate these enhancements through extensive ablation studies and show that GraFT consistently surpasses established multimodal ReID benchmarks. Additionally, aiming for deployment versatility, we've integrated neural network pruning into GraFT, offering a balance between model size and performance.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Multi-Modal Object Re-Identification with Prompt-S6 and Semantic-Aware Knowledge Guidance

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Prompt-S6 plus semantic token pruning and progressive tri-modal fusion improves multi-spectral object ReID accuracy and efficiency on four benchmarks.

  2. ICPL-ReID: Identity-Conditional Prompt Learning for Multi-Spectral Object Re-Identification

    cs.CV 2025-05 conditional novelty 6.0 of 10

    An identity-conditioned, online prompt learning framework with low-rank adapters sets new state-of-the-art results on five multi-spectral person and vehicle re-identification benchmarks.

  3. MambaPro: Multi-Modal Object Re-Identification with Mamba Aggregation and Synergistic Prompt

    cs.CV 2024-12 conditional novelty 6.0 of 10

    MambaPro combines CLIP with a parallel adapter, synergistic residual prompts, and Mamba aggregation to achieve state-of-the-art mAP on RGBNT201, RGBNT100, and MSVR310.

  4. DeMo: Decoupled Feature-Based Mixture of Experts for Multi-Modal Object Re-Identification

    cs.CV 2024-12 conditional novelty 6.0 of 10

    DeMo improves multi-modal object re-identification by decoupling RGB, NIR, and TIR features into seven attention-derived streams and weighting them with an attention-triggered mixture of experts.

  5. Multi-Modal Object Re-Identification with Dual Semantic Guidance and Global-Local Mutual Modulation

    cs.CV 2026-07 conditional novelty 4.0 of 10

    A dual-semantic (text + soft mask) global-local mutual modulation framework reports SOTA mAP/Rank-1 on RGBNT201, RGBNT100, and MSVR310.

Pith tools