Pith. sign in

REVIEW 7 cited by

Semantically Self-Aligned Network for Text-to-Image Part-aware Person Re-identification

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2107.12666 v2 pith:OKMK2XN2 submitted 2021-07-27 cs.CV

classification cs.CV
keywords ssantext-to-imagetextualdescriptionsnetworkpersonreidsemantically
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Text-to-image person re-identification (ReID) aims to search for images containing a person of interest using textual descriptions. However, due to the significant modality gap and the large intra-class variance in textual descriptions, text-to-image ReID remains a challenging problem. Accordingly, in this paper, we propose a Semantically Self-Aligned Network (SSAN) to handle the above problems. First, we propose a novel method that automatically extracts semantically aligned part-level features from the two modalities. Second, we design a multi-view non-local network that captures the relationships between body parts, thereby establishing better correspondences between body parts and noun phrases. Third, we introduce a Compound Ranking (CR) loss that makes use of textual descriptions for other images of the same identity to provide extra supervision, thereby effectively reducing the intra-class variance in textual features. Finally, to expedite future research in text-to-image ReID, we build a new database named ICFG-PEDES. Extensive experiments demonstrate that SSAN outperforms state-of-the-art approaches by significant margins. Both the new ICFG-PEDES database and the SSAN code are available at https://github.com/zifyloo/SSAN.

Discussion (0). Sign in to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Dual-Granularity Cross-Modal Identity Association for Weakly-Supervised Text-to-Person Image Matching

    cs.CV 2025-07 conditional novelty 7.0 of 10

    A dual-granularity identity association mechanism with dynamic confidence weighting improves weakly supervised text-to-person matching, achieving 73.06% Rank-1 on CUHK-PEDES.

  2. Achieving Text-based Person Retrieval with Any Granularity

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A multi-grained dataset, evaluation benchmark, and CMAM model are proposed for text-based person retrieval across five levels of query granularity, with CMAM outperforming existing methods on the new benchmark.

  3. Gradient-Attention Guided Dual-Masking Synergetic Framework for Robust Text-based Person Retrieval

    cs.CV 2025-09 conditional novelty 6.0 of 10

    GA-DMS with the WebPerson dataset sets new state-of-the-art Rank-1 accuracy on CUHK-PEDES, ICFG-PEDES, and RSTPReid.

  4. Dynamic Pattern Alignment Learning for Pretraining Lightweight Human-Centric Vision Models

    cs.CV 2025-08 conditional novelty 6.0 of 10

    Distilling three pattern-specific alignments from a large human-centric teacher yields a 5M-parameter student that approaches teacher-level generalization on many downstream tasks.

  5. Human-centered Interactive Learning via MLLMs for Text-to-Image Person Re-identification

    cs.LG 2025-05 conditional novelty 6.0 of 10

    A test-time MLLM interaction module and a text reorganization augmentation improve text-to-image person re-identification ranking across four benchmarks.

  6. Blurring Modal Boundaries: A Unified Survey from Single- to Multi-Modal Person Re-ldentification

    cs.CV 2026-07 conditional novelty 5.0 of 10

    A unified taxonomy and survey of single- to multi-modal person ReID, plus a Transformer-based VI-ReID baseline that is solid but not state-of-the-art.

  7. Adapting Vision-Language Models Without Labels: A Comprehensive Survey

    cs.LG 2025-08 conditional novelty 5.0 of 10

    A survey that organizes unsupervised vision-language model adaptation by unlabeled-data availability into four paradigms: data-free transfer, domain transfer, episodic test-time, and online test-time adaptation.

Pith tools